
Latent Space: The AI Engineer Podcast · Latent.Space
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light" — Nader Khalil (Brev), Kyle Kranen (Dynamo)
·1 hr 24 min·5 clips
Dynamo is NVIDIA’s data-center-scale inference engine for scaling out beyond a single model replica.
As heard by us
Agent serving gets real when security, latency, and cost all have to be balanced.
The discussion treats agent inference as an operating problem, not just a model problem. It argues that when an agent can reach files, the internet, and custom code, the practical move is to keep any two of the three and set clear enforcement points.
Why you'd press play
How NVIDIA thinks about agent security when files, internet, and code collide.
Listen to the show on