C. Braga, Manuel A. Serrano, E. Fernández-Medina
2026.8.1KNOWLEDGE-BASED SYSTEMS
Abstract
AI systems operating in complex, high-frequency environments often make critical decisions such as generating alerts, prioritizing events or assigning severity scores , under conditions of partial ground truth , uneven group visibility and implicit, context-dependent logic. In such settings, it becomes difficult to assess whether behavior is consistent across sources, categories or technical modules. This challenge is particularly relevant in domains such as cybersecurity, where automated decisions directly influence threat prioritization, resource allocation and system-level risk exposure. Fairness metrics have been extensively studied in machine learning, yet most approaches assume supervised classification tasks with complete labels and explicit demographic groups. These assumptions often break down in operational AI systems, where decisions are decentralized, labels are partial or delayed, and group-defining attributes are technical , such as protocol type, sensor origin or detection source. As a result, classical fairness metrics often fail to detect or characterize structural disparities in real-world, partially supervised deployments. To address this gap, we propose a structured six-phase methodology for adapting fairness metrics, specifically independence , separation and sufficiency , to decision systems operating under dynamic and context-sensitive conditions. The framework is applied to the domain of cybersecurity, where the analysis of system behavior is challenged by alert inflation, detection imbalance and score inconsistency. The metrics are characterized using four synthetic scenarios, crafted to reproduce typical diagnostic conditions and to verify the selective sensitivity of each metric, and a real-world intrusion detection dataset (UNSW-NB15), in which a trained Random Forest classifier is evaluated via 5-fold stratified cross-validation to reflect operational conditions. Results show that, even under a high-performing classifier (global F1 = 0.98), the proposed metrics reveal protocol-level disparities not captured by aggregate performance measures, and remain stable under consistency-preserving conditions. Statistical testing confirms that the observed disparities are highly significant ( p < 1 0 − 300 ), with effect size analysis confirming their practical magnitude, and comparative analysis with standard fairness metrics reveals the limited diagnostic resolution of these classical indicators in operational settings . The resulting metrics are mutually compatible and respond independently across decision dimensions without conflict. Their interpretability and traceability support structured diagnostic analysis of system behavior. We formally derive their theoretical properties , and discuss how they complement the diagnostic objectives promoted by governance frameworks such as the EU AI Act, the NIST AI Risk Management Framework and ISO/IEC 42001. These results position the resulting metrics as diagnostic instruments for analyzing system-level behavior in environments characterized by partial supervision and evolving behavior .
Citation format
BRAGA, C.; SERRANO, Manuel A.; FERNÁNDEZ-MEDINA, E. Operational fairness diagnostics for AI-based detection systems: Metrics and methodology. KNOWLEDGE-BASED SYSTEMS, 2026.