Latent Space: The AI Engineer Podcast · Latent.Space

Artificial Analysis: The Independent LLM Analysis House — with George Cameron and Micah Hill-Smith

·1 hr 18 min·5 clips
Artificial Analysis uses a mystery shopper policy to secretly test lab endpoints, ensuring no one can manipulate their independent benchmarks.
George and Micah, co-founders of Artificial Analysis, discuss the origins and evolution of their independent AI benchmarking platform. They cover challenges in model evaluation, new metrics like the Omniscience Index for hallucination, and trends in AI cost and performance, emphasizing the importance of transparency and industry impact.

As heard by us

A sharp look at how AI benchmarking becomes something developers and companies can actually use.

Artificial Analysis comes across as an independent third-party site for comparing models and hosts, and the value is clear: it turns raw data into a practical way for developers and companies to weigh quality against throughput.

Read the full review in PlayNext →

Why you'd press play

You want the backstory behind the benchmark site people keep citing: a free, independent third-party comparison that breaks out quality versus throughput by model and hosting provider.

Read the full recommendation in PlayNext →
Listen to the show on