Utilizing Tech - The Podcast Series about New and Emerging Technologies · Tech Field Day - Part of The Futurum Group

07x02: Building an AI Training Data Pipeline with VAST Data

·31 min·3 clips
Kartik explains how enterprises struggle with tens of petabytes of data to wrangle for AI training.
This episode examines the complex data pipeline required for generative AI, moving beyond a narrow focus on GPU training. Host Stephen Foskett and co-host Janice Narowski from Solidigm are joined by Kartik Subramanian, a data platform expert from Vast Data with a background in particle physics. They discuss the full lifecycle of AI data, from initial collection to final inference. The conversation begins by challenging the common overemphasis on GPU compute, arguing that data preparation is a much larger and more difficult challenge. Kartik notes that while GPT-3 was trained on roughly 500GB of tokens, that data was distilled from petabytes of raw internet information. He explains that enterprises now face a massive data wrangling problem, needing to incorporate tens or hundreds of petabytes of internal data into their AI pipelines. Vast Data approaches this by offering a unified data platform that consolidates storage, exposes data in tabular formats for analysis, and aims to eliminate unnecessary data movement. A key insight is that the AI data pipeline is data-intensive at every stage, not just during model training. Kartik predicts a $2.8 trillion infrastructure transformation over five years as enterprises aggregate and understand their data to become "AI ready." He differentiates between generative AI and traditional enterprise analytics, where 95% of work still uses classic machine learning on structured data. The discussion covers how High-Performance Computing (HPC) centers are also transforming, blending traditional simulation workloads with new, random I/O-intensive AI workloads. The episode highlights the critical role of all-flash storage for predictable performance, especially during model checkpointing, and notes the economic and power efficiency advantages of high-capacity solid-state drives. Kartik shares that Vast Data works with AI-focused cloud service providers like CoreWeave and Lambda, which demand robust security, governance, and multi-tenant data separation. The EU AI Act is mentioned as a driver for new technical challenges, mandating the long-term preservation of all queries and data used in high-risk inference. The tone is educational and conversational, featuring a deep technical discussion between industry practitioners. The hosts guide the expert guest through a structured exploration of infrastructure challenges and market trends. This episode is ideal for IT infrastructure architects, data engineers, and anyone involved in planning or building enterprise AI systems. Listeners solely interested in AI model theory or consumer applications might find the deep infrastructure focus less relevant.
Listen to the show on