MAFT: Efficient Model-Agnostic Fairness Testing for Deep Neural Networks via Zero-Order Gradient Search
Zhaohui Wang, Min Zhang, Jingran Yang, Bojie Shao, Min Zhang
Abstract
Deep neural networks (DNNs) have shown powerful performance in various applications and are increasingly being used in decisionmaking systems. However, concerns about fairness in DNNs always persist. Some efficient white-box fairness testing methods about individual fairness have been proposed. Nevertheless, the development of black-box methods has stagnated, and the performance of existing methods is far behind that of white-box methods. In this paper, we propose a novel black-box individual fairness testing method called Model-Agnostic Fairness Testing (MAFT). By leveraging MAFT, practitioners can effectively identify and address discrimination in DL models, regardless of the specific algorithm or architecture employed. Our approach adopts lightweight procedures such as gradient estimation and attribute perturbation rather than non-trivial procedures like symbol execution, rendering it significantly more scalable and applicable than existing methods. We demonstrate that MAFT achieves the same effectiveness as stateof-the-art white-box methods whilst improving the applicability to large-scale networks. Compared to existing black-box approaches, our approach demonstrates distinguished performance in discovering fairness violations w.r.t effectiveness (∼ 14.69×) and efficiency (∼ 32.58×). CCS CONCEPTS • Computing methodologies → Artificial intelligence; • Software and its engineering → Software testing and debugging.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 69e5be33-7255-4585-9862-7516e7293df8Cited by top-tier papers4
- Fairness Mediator: Neutralize Stereotype Associations to Mitigate Bias in Large Language ModelsYisong Xiao, Aishan Liu, Siyuan Liang, Xianglong Liu et al.ISSTA 2025 · 2 citations
- Fairness Invariants: A Relational Approach to Explaining and Mitigating Fairness BugsRanit Debnath Akash, Ashish Kumar, Gang Tan, Saeid Tizpaz-NiariISSTA 2026
- Perturbation Effects on Robustness and Individual FairnessXuran Li, Hao Xue, Peng Wu, Xingjun Ma et al.KDD 2026
- Provable Fairness Repair for Deep Neural NetworksJianan Ma, Jingyi Wang, Qi Xuan, Zhen WangASE 2025
Builds on6
- Fairness-aware News Recommendation with Decomposed Adversarial LearningChuhan Wu, Fangzhao Wu, Xiting Wang, Yongfeng Huang et al.AAAI 2021 · 176 citations
- White-box fairness testing through adversarial samplingPeixin Zhang, Jingyi Wang, Jun Sun, Guoliang Dong et al.ICSE 2020 · 127 citations
- NeuronFair: Interpretable White-Box Fairness Testing through Biased Neuron IdentificationHaibin Zheng, Zhiqing Chen, Tianyu Du, Xuhong Zhang et al.ICSE 2022 · 58 citations
- Efficient white-box fairness testing through gradient searchLingfeng Zhang, Yueling Zhang, Min ZhangISSTA 2021 · 51 citations
- Explanation-Guided Fairness Testing through Genetic AlgorithmMing Fan, Wenying Wei, Wuxia Jin, Zijiang Yang et al.ICSE 2022 · 45 citations
Related papers
- Approximation-guided Fairness Testing through Discriminatory Space AnalysisZhenjiang Zhao, Takahisa Toda, Takashi KitamuraASE 2024
- Dissecting Global Search: A Simple Yet Effective Method to Boost Individual Discrimination Testing and RepairLili Quan, Tianlin Li, Xiaofei Xie, Zhenpeng Chen et al.ICSE 2025 · 2 citations
- RULER: discriminative and iterative adversarial training for deep neural network fairnessGuanhong Tao, Weisong Sun, Tingxu Han, Chunrong Fang et al.FSE 2022 · 29 citations
- Fairquant: Certifying and Quantifying Fairness of Deep Neural NetworksBrian Hyeongseok Kim, Jingbo Wang, Chao WangICSE 2025 · 6 citations
- Fairify: Fairness Verification of Neural NetworksSumon Biswas, Hridesh RajanICSE 2023 · 27 citations
