Latent Space: The AI Engineer Podcast · Latent.Space

Artificial Analysis: Independent LLM Evals as a Service — with George Cameron and Micah-Hill Smith

·1 hr 18 min·7 clips
AI labs sometimes give special endpoints for benchmarking that might not match what users actually get.
This episode features George Cameron and Micah-Hill Smith, founders of the independent AI benchmarking platform Artificial Analysis. They discuss their journey from a side project to a company with over 20 employees, serving enterprise clients and AI builders. The conversation explores the necessity of independent evaluation in a rapidly evolving AI model landscape. Artificial Analysis began in 2023 as a free public website comparing model quality, throughput, and cost trade-offs. The founders initially funded the benchmarking costs personally, which were relatively low at the time. A key early challenge was ensuring evaluation consistency, as different labs used varied prompting techniques that could skew results. They developed a "mystery shopper" policy to verify that models perform identically on private versus public endpoints. The company later joined AI Grant, gaining mentorship and aligning with other frontier-building companies. The platform's flagship Artificial Analysis Intelligence Index synthesizes data from ten different evaluation datasets into a single score. The index has evolved through several versions to avoid saturation, incorporating agentic capabilities and long-context reasoning. A notable new evaluation is the Omniscience Index, designed to measure factual knowledge and hallucination rates by penalizing incorrect answers. The founders observed that Claude models exhibited the lowest hallucination rates in this test, while Gemini 3 Pro showed a significant accuracy leap. Interestingly, the Omniscience Index's accuracy metric closely correlates with total model parameter count, providing indirect clues about undisclosed model sizes. The founders note that general intelligence scores do not strongly correlate with a model's tendency to avoid hallucinations. They also highlight the Critical Point evaluation, a physics problem set where even top models score only 9%, illustrating areas where controlled "hallucination" can be a creative feature. The tone is conversational and deeply technical, reflecting the hosts' and guests' shared expertise in AI infrastructure. The style is educational, walking through specific charts, methodologies, and industry anecdotes. This episode is ideal for AI engineers, product managers, and enterprise decision-makers who rely on model performance data. Listeners seeking high-level business trends without technical depth might find the detailed evaluation discussions less engaging.

As heard by us

Independent AI benchmarking turns into a business story, then into a debate about what model scores really measure.

Artificial Analysis reads like an independent third-party comparison site for model quality and throughput, with model and hosting provider breakdowns that give it a practical, infrastructural feel.

Read the full review in PlayNext →

Why you'd press play

You get the business model, the eval math, and the part where everyone argues with the numbers.

Read the full recommendation in PlayNext →
Listen to the show on