Latent Space: The AI Engineer Podcast · Latent.Space

NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light" — Nader Khalil (Brev), Kyle Kranen (Dynamo)

·1 hr 24 min·5 clips
Dynamo is NVIDIA’s data-center-scale inference engine for scaling out beyond a single model replica.
The cold open starts with a clean constraint: agents can touch files, reach the internet, and now write and run custom code.

As heard by us

Agent serving gets real when security, latency, and cost all have to be balanced.

The discussion treats agent inference as an operating problem, not just a model problem. It argues that when an agent can reach files, the internet, and custom code, the practical move is to keep any two of the three and set clear enforcement points.

Read the full review in PlayNext →

Why you'd press play

How NVIDIA thinks about agent security when files, internet, and code collide.

Read the full recommendation in PlayNext →
Listen to the show on