Talk Python To Me · Michael Kennedy

Python apps with LLM building blocks

November 30, 2025·1 hr 17 min·6 clips
Vincent Wormerdam explains how to treat LLMs as just another API in your Python app with clear boundaries and good monitoring.
Episode 528 starts with boundaries. Michael sets up LLMs as another API surface inside Python applications, with edges you can see and debug instead of vague magic. The conversation stays practical. The excerpt points to small focused endpoints, explicit wrappers, monitoring, caching, response inspection, and the question of where an LLM call is worth the trouble. Then it turns into a real caching problem. Michael describes a site where image URLs carry cache-busting IDs. That works locally, but gets awkward once resources live in S3-style blob storage. The issue is familiar: after an app restarts, it may not know whether a remote file changed unless it downloads much more than it wants. The discussion moves into the middle ground between a full CDN and raw blob fetches, where a local cache can take the sting out. Michael picks up the shape of it. The site downloads an external resource once, computes a hash, stores it in a disk cache, and reuses it until something changes. That helps avoid stale resources while making remote assets feel quick, even when the app does not own the source of truth. The TTL question is where the design gets honest. Michael separates assets managed through admin functions from parsed description content that can refresh on a daily-style cadence. Admin changes can invalidate directly. Other derived resources can use time-to-live behavior, because the upstream trigger is not always under the app's control. The best part is the tone. This is not architecture theater. It is two Python developers walking a real annoyance toward something simple enough to trust.

As heard by us

A practical Python conversation about treating LLM calls as manageable, observable parts of an application.

Python apps with LLM building blocks treats model calls as ordinary production dependencies: API calls with boundaries, monitoring, caches, and a real question of whether they belong in the architecture at all.

Read the full review in PlayNext →

Why you'd press play

Treat LLMs like a normal API, not a magic layer.

Read the full recommendation in PlayNext →
Listen to the show on