This is Fine! A podcast about resilience engineering and software · Colette Alexander and Clint Byrum

What's the ROI on Reliability and Resilience Work?

·58 min·5 clips
How do you define the ROI of reliability work when you can't prove a negative?
This episode features hosts Clint and Colette exploring how to measure the return on investment for reliability and resilience engineering work. They respond to a listener question from the Resilience and Software Foundation Slack about defining and justifying the ROI of such efforts. Clint and Colette are practitioners who regularly discuss the practical challenges of building resilient systems. The central question asks how to surface the ROI of reliability features to win prioritization battles against other business demands. Clint references Dave Woods' idea that resilience is about "what you can do" rather than "what you have," which is harder to measure. They discuss Jens Rasmussen's system model, which describes economic, physical, and safety boundaries that constantly pressure organizations. Colette notes that reliability work often involves preventing negatives, which is inherently difficult to prove or quantify. A key insight is that significant cultural change often requires a sufficient amount of externalized pain, such as customer complaints or revenue risk, to create a window for investment. The hosts emphasize the necessity of involving product managers to define what reliability means for customer value and to segment basic expectations from nice-to-have features. They compare foundational reliability work to code review—a practice whose ROI is rarely calculated but is accepted as a cultural necessity. Clint expresses frustration with the lack of satisfying metrics, suggesting lagging indicators like reduced engineer turnover or exhaustion from on-call shifts. They humorously reference a "sentiment button" for qualitative vibes as a potential gauge. The conversation acknowledges that during periods of high labor availability, leaders might deprioritize engineer sentiment in favor of immediate economic pressures, despite the long-term risk of critical brain drain. The tone is conversational and reflective, blending personal anecdotes with professional frameworks from resilience engineering. The style is educational, grounded in the hosts' shared experiences navigating technical and organizational trade-offs. Listeners interested in the organizational politics of engineering and practical strategies for advocating for foundational work will appreciate this episode. Those seeking a definitive, quantitative answer for calculating ROI might find the discussion unsatisfying, as it concludes that cultural integration is often the goal.

As heard by us

A clear-eyed conversation about making hidden reliability pain visible enough for the rest of the organization to act on.

It treats reliability as an organizational problem, not a technical side note. The strongest stretch is where it moves from that familiar pressure point to practical work: SLOs, synthetic tests, and the awkward but necessary job of sitting down with product to define what…

Read the full review in PlayNext →

Why you'd press play

When reliability work needs the pain to be visible, this episode shows how to externalize it.

Read the full recommendation in PlayNext →
Listen to the show on