Latent Space: The AI Engineer Podcast · Latent.Space

🔬 Automating Science: World Models, Scientific Taste, Agent Loops — Andrew White

·1 hr 14 min·5 clips
MD had the vibe of something that should eventually win. Andrew remembers D.E. Shaw Research as the serious version of that bet: custom silicon, custom clusters, and the belief that molecular dynamics could fold proteins if you pushed hard enough. Then AlphaFold arrived from another angle. The surprising part was that protein folding worked on ordinary tools like Google Colab, a GPU, or a desktop. That changes the story. A problem that seemed to need special machines suddenly looked more like a problem of model choice and representation. The interview keeps pressing on what changed, and the answer is messier than better data plus bigger compute. Scientific data already carries judgment inside it. Experts can look at the same dataset and walk away with different reads. The host ties that to Heather Kulik finding raw data in papers that did not match the papers' own conclusions. Andrew treats that as the real problem, not a side note. BIGS Bench becomes the pressure test. It is their bioinformatics benchmark, and some frontier LLM system cards mention it as one of the tests they use. Scores now sit around 60 to 70 correctness, which is a strange place to be. If humans agree on only about 70 percent of the analysis, the benchmark is measuring field disagreement along with model mistakes. That makes the product question awkward in a useful way. Companies are not only asking whether an agent is correct. They want to know whether it follows their lab's habits, including when to impute data and when to refuse. The agent has to learn the local taste. Objective analysis starts to look like choices, norms, and domain instincts that nobody has fully agreed how to score. Then medicinal chemistry shows why the hard parts stay hard. Andrew calls it a modern dark art in the AI-resistant zone, where nomenclature, reagents, and tiny data format choices can derail the model. Alphabetical sorting can even fool it. If the training data uses one ordering convention and the test data uses another, the model may learn the formatting trick instead of the chemistry. The close feels very Latent Space. A dense technical thread ends up in hiring, infrastructure, Edison Scientific, and the blunt reality that science agents live or die on benchmarks, messy data, and dull-looking details that run more of the game than anyone wants to admit.

As heard by us

AI in science gets real when the data stops behaving.

Andrew White's episode tracks a simple but useful shift: AI is no longer just an idea floating above science, it is running into the daily friction of real work.

Read the full review in PlayNext →

Why you'd press play

When protein folding moves from a hardware moonshot to a desktop demo, the field changes fast.

Read the full recommendation in PlayNext →
Listen to the show on