Provable Fairness Repair for Deep Neural Networks
Jianan Ma, Jingyi Wang, Qi Xuan, Zhen Wang
Abstract
Deep neural networks (DNNs) are suffering from ethical issues such as individual discrimination. In response, extensive NN repair techniques have been developed to adjust models and mitigate such undesired behaviors. However, existing fairness repair methods are typically data-centric, which often lack provable guarantees and generalization to unseen samples. To overcome these limitations, we propose PROF, a novel fairness repair framework with provable guarantees. The key intuition of PROF is to leverage interval bound propagation (a widely used NN verification technique) to soundly capture model outputs over the whole set around a biased sample x. The derived bounds are utilized to guide fairness repair which encourages the model to produce consistent outputs on . Specifically, we integrate fairness constraints and model modifications into a unified constraint-solving formulation, which can be transformed to a Mixed-Integer Linear Programming (MILP) problem solvable by off-the-shelf solvers. The solution to the MILP problem effectively induces a repaired model with guaranteed fairness over the whole set . We evaluate PROF on four widely used benchmark datasets and demonstrate that it achieves provable fairness repair, with generalization of up to 95.93% on full datasets and 93.16% on the entire input space. Notably, PROF can be easily configured to support multiple sensitive attributes and more practical fairness definitions, while providing provable repair guarantees and delivering around 90% fairness improvement. Our code is available in this .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2d5ebab-dd64-4123-867a-3d27d996a290Builds on24
- AI2: Safety and Robustness Certification of Neural Networks with Abstract InterpretationTimon Gehr, Matthew Mirman, Dana Drachsler-Cohen, Petar Tsankov et al.S&P 2018 · 987 citations
- Formal Security Analysis of Neural Networks using Symbolic IntervalsShiqi Wang, Kexin Pei, Justin Whitehouse, Junfeng Yang et al.USENIX Security 2018 · 523 citations
- Automatic Perturbation Analysis for Scalable Certified Robustness and BeyondKaidi Xu, Zhouxing Shi, Huan Zhang, Yihan Wang et al.NeurIPS 2020 · 415 citations
- White-box fairness testing through adversarial samplingPeixin Zhang, Jingyi Wang, Jun Sun, Guoliang Dong et al.ICSE 2020 · 127 citations
- Causality-Based Neural Network RepairBing Sun, Jun Sun, Long H. Pham, Tie ShiICSE 2022 · 69 citations
Related papers
- RULER: discriminative and iterative adversarial training for deep neural network fairnessGuanhong Tao, Weisong Sun, Tingxu Han, Chunrong Fang et al.FSE 2022 · 29 citations
- Patch Synthesis for Property Repair of Deep Neural NetworksZhiming Chi, Jianan Ma, Pengfei Yang, Cheng-Chao Huang et al.ICSE 2025 · 2 citations
- Fairquant: Certifying and Quantifying Fairness of Deep Neural NetworksBrian Hyeongseok Kim, Jingbo Wang, Chao WangICSE 2025 · 6 citations
- REGLO: Provable Neural Network Repair for Global Robustness PropertiesFeisi Fu, Zhilu Wang, Weichao Zhou, Yixuan Wang et al.AAAI 2024 · 11 citations
- NeuronFair: Interpretable White-Box Fairness Testing through Biased Neuron IdentificationHaibin Zheng, Zhiqing Chen, Tianyu Du, Xuhong Zhang et al.ICSE 2022 · 58 citations
