Software Engineering Institute (SEI) Podcast Series · Members of Technical Staff at the Software Engineering Institute

Improving Machine Learning Test and Evaluation with MLTE

·29 min·1 clip
Kate names two failure modes: missing system requirements and weak communication between model developers and stakeholders.
1. The SEI Podcast Series episode "Improving Machine Learning Test and Evaluation with MLTE" focuses on the MELT tool for machine learning tests and evaluation. 2. Grace Lewis hosts the discussion as principal researcher and lead of the SEI's Tactical and AI-Enabled Systems Initiative, joined by Sebastiana Chavarria, Alex Durr, and Army data scientist Kate Maffey. 3. The episode asks what MELT is for, and the answer is a system-level process and tool for ML testing that captures requirements before integration. 4. Sebastiana says MELT came from seeing ML-enabled components increase across systems while expectations stayed mismatched between model developers and system developers. 5. She says models can work in isolation and still fail when they are deployed into production systems. 6. MELT starts with quality attributes and a cyclical discussion between stakeholders and model developers during design time. 7. Alex describes the negotiation card as the artifact that records the discussion and identifies pain points such as CPU availability, memory availability, and operating environment. 8. He says MELT then helps teams define test cases through the MELT library and test catalog. 9. The test catalog includes examples such as robustness to blur, and organizations can add their own tests to an internal instance. 10. The MELT library is designed to run tests and generate a PDF report showing requirements, test choices, and results. 11. Kate says ML models fail when people do not understand system requirements, such as building an image detector without knowing it will run on a handheld device. 12. She says weak communication can leave one domain expert holding critical operational knowledge while the model developer writes code without it. 13. Grace connects that problem to data scientists who default to accuracy testing instead of testing system behavior. 14. Kate says stakeholders should include domain experts, system owners, and users because each group defines success differently. 15. She says the negotiation card matters because it creates formal touch points that teams rarely organize on their own. 16. She also says MELT revisits the negotiation after results come back, so teams can redefine impossible requirements like speed or accuracy. 17. Sebastiana says MELT 1.0 is open source on GitHub and includes documentation, tutorials, demos, and linked papers. 18. Alex says the Continuum project is a three-year LTP focused on improving MELT and transitioning it to the DoD and integrated T&E. 19. The episode fits listeners who work on ML systems, test and evaluation, or DoD software engineering. 20. It is less useful for listeners who want a product demo without process discussion.

As heard by us

A practical look at how MELT turns ML test and evaluation into a process teams can work with.

Machine learning test and evaluation sits at the center of this episode, and its real value is in making MELT feel usable at design time. Grace Lewis keeps the focus on deployment realities, while Alex Durr, Sebastiana Chavarria, and Kate Maffey turn concerns like CPU, memory,…

Read the full review in PlayNext →

Why you'd press play

Listen if you want ML testing turned into a process, not just a slogan.

Read the full recommendation in PlayNext →
Listen to the show on