Latent Space: The AI Engineer Podcast · Latent.Space

The First Mechanistic Interpretability Frontier Lab — Myra Deng & Mark Bissell of Goodfire AI

·1 hr 8 min·5 clips
Watch as Mark steers a trillion-parameter AI model to suddenly start talking in Gen Z slang during a live demo.
Mark and Myra from Goodfire appear on Latent Space alongside special co-host Vibhu and Mochi the mechanistic interpretability dog to discuss the startup's $150 million Series B at a $1.25 billion valuation. Goodfire describes itself as an AI research lab that uses interpretability to understand, learn from, and design AI models, with a belief that interpretability will unlock the next frontier of safe and powerful AI. Mark joined as one of Goodfire's first ten employees after three years at Palantir as a forward-deployed engineer on healthcare teams; Myra came from Two Sigma and now leads product. Vibhu opens with the question of what interpretability actually means, and the Goodfire team describes it broadly as understanding what is happening inside a neural network at the level of features, circuits, and representations rather than just observing input-output behavior. The discussion traces the field's evolution from Anthropic's early toy models showing superposition, where models represent more concepts than they have dimensions, to sparse autoencoders that can identify interpretable features in large production models. Goodfire's Ember API provides programmatic access to model internals, enabling developers to identify what features activate for specific inputs, suppress or amplify those features, and observe how the model behaves as a result. The team discusses production use cases including life sciences research, where interpretability helps scientists understand what a model is paying attention to when making predictions, and healthcare systems where regulatory requirements demand explainability that black-box evaluation cannot provide. Mark describes the transition from research demos to real products as requiring a shift from studying whether a model has a banana feature to asking what sequence of features fires when a model makes a clinical recommendation. The fundraise announcement is framed as validation that interpretability has moved from a niche Anthropic safety research concern into a commercial category with real enterprise demand. The hosts discuss the Goodfire team structure, which is deliberately generalist, with members switching between research, engineering, and product depending on what the company needs. The conversation ends on the open question of whether interpretability will eventually be mandatory for regulated AI deployments, which Goodfire sees as the most likely driver of long-term enterprise adoption.
Listen to the show on