Lynn Vonderhaar, Timothy Elvira, Omar Ochoa
tlooto Summary
Current literature on provenance for ML is aggregates and provides a method of measurement for otherwise unquantifiable NFRs and analyzes trends in how provenance can decompose various ML NFRs and future directions for the field.
Abstract
As Machine Learning (ML) becomes ever more ubiquitous, it is critical to increase the rigor of design and testing. In traditional software, this is done using the Requirements Engineering (RE) process, but RE looks different for ML because it is data centric. There are different Non-Functional Requirements (NFRs) and standard testing techniques do not apply. While there are standards for verifying NFRs in traditional software, there is no standard measurement for ML NFRs, e.g., how is a model verified to meet an explainability NFR? In traditional software, NFRs are decomposed into Functional Requirements (FRs), but without clear measurements for ML NFRs, their decomposition into FRs is nearly impossible. However, recently, research has shown that provenance can help improve model transparency and reproducibility. This work builds on such literature and suggests provenance as a lower-level NFR to connect high-level NFRs, e.g., explainability and transparency, and FRs, thereby enabling concrete model verification based on requirement specifications. This work examines types of ML provenance and their use in decomposing model NFRs into verifiable FRs, thereby better aligning ML development with RE and increasing the rigor of ML testing. This paper aggregates current literature on provenance for ML and provides a method of measurement for otherwise unquantifiable NFRs. This work also analyzes trends in how provenance can decompose various ML NFRs and future directions for the field.
Citation format
VONDERHAAR, Lynn; ELVIRA, Timothy; OCHOA, Omar. Provenance as a machine learning non-functional requirement: Trends and future directions. International Journal of Semantic Computing, 2026, 20(01): 135–156.