Latent Imitator: Generating Natural Individual Discriminatory Instances for Black-Box Fairness Testing
Yisong Xiao, Aishan Liu, Tianlin Li, Xianglong Liu
摘要
Machine learning (ML) systems have achieved remarkable performance across a wide area of applications. However, they frequently exhibit unfair behaviors in sensitive application domains (e.g., employment and loan), raising severe fairness concerns. To evaluate and test fairness, engineers often generate individual discriminatory instances to expose unfair behaviors before model deployment. However, existing baselines ignore the naturalness of generation and produce instances that deviate from the real data distribution, which may fail to reveal the actual model fairness since these unnatural discriminatory instances are unlikely to appear in practice. To address the problem, this paper proposes a framework named Latent Imitator (LIMI) to generate more natural individual discriminatory instances with the help of a generative adversarial network (GAN), where we imitate the decision boundary of the target model in the semantic latent space of GAN and further samples latent instances on it. Specifically, we first derive a surrogate linear boundary to coarsely approximate the decision boundary of the target model, which reflects the nature of the original data distribution. Subsequently, to obtain more natural instances, we manipulate random latent vectors to the surrogate boundary with a one-step movement, and further conduct vector calculation to probe two potential discriminatory candidates that may be more closely located in the real decision boundary. Extensive experiments on various datasets demonstrate that our LIMI outperforms other baselines largely in effectiveness (×9.42 instances), efficiency (×8.71 speeds), and naturalness (+19.65%) on average. In addition, we empirically demonstrate that retraining on test samples generated by our approach can lead to improvements in both individual fairness (45.67% on 𝐼 𝐹 𝑟 and 32.81% on 𝐼 𝐹 𝑜 ) and group fairness (9.86% on 𝑆𝑃𝐷 and 28.38% on 𝐴𝑂𝐷). Our codes can be found on our website [7] .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- FAIRER: Fairness as Decision Rationale AlignmentTianlin Li, Qing Guo, Aishan Liu, Mengnan Du 等ICML 2023 · 被引用 20 次
- Detoxifying Large Language Models via Autoregressive Reward Guided Representation EditingYisong Xiao, Aishan Liu, Siyuan Liang, Zonghao Ying 等NeurIPS 2025 · 被引用 12 次
- RUNNER: Responsible UNfair NEuron Repair for Enhancing Deep Neural Network FairnessTianlin Li, Yue Cao, Jian Zhang, Shiqian Zhao 等ICSE 2024 · 被引用 11 次
- GenderCARE: A Comprehensive Framework for Assessing and Reducing Gender Bias in Large Language ModelsKunsheng Tang, Wenbo Zhou, Jie Zhang, Aishan Liu 等CCS 2024 · 被引用 7 次
- MAFT: Efficient Model-Agnostic Fairness Testing for Deep Neural Networks via Zero-Order Gradient SearchZhaohui Wang, Min Zhang, Jingran Yang, Bojie Shao 等ICSE 2024 · 被引用 6 次
它引用的顶会 Paper11
- Bias in machine learning software: why? how? what to do?Joymallya Chakraborty, Suvodeep Majumder, Tim MenziesFSE 2021 · 被引用 186 次
- Model-based exploration of the frontier of behaviours for deep learning system testingVincenzo Riccio, Paolo TonellaFSE 2020 · 被引用 134 次
- Fairway: a way to build fair ML softwareJoymallya Chakraborty, Suvodeep Majumder, Zhe Yu, Tim MenziesFSE 2020 · 被引用 131 次
- White-box fairness testing through adversarial samplingPeixin Zhang, Jingyi Wang, Jun Sun, Guoliang Dong 等ICSE 2020 · 被引用 127 次
- Fairea: a model behaviour mutation approach to benchmarking bias mitigation methodsMax Hort, Jie M. Zhang, Federica Sarro, Mark HarmanFSE 2021 · 被引用 75 次
相关 Paper
- Efficient white-box fairness testing through gradient searchLingfeng Zhang, Yueling Zhang, Min ZhangISSTA 2021 · 被引用 51 次
- NeuronFair: Interpretable White-Box Fairness Testing through Biased Neuron IdentificationHaibin Zheng, Zhiqing Chen, Tianyu Du, Xuhong Zhang 等ICSE 2022 · 被引用 58 次
- Dissecting Global Search: A Simple Yet Effective Method to Boost Individual Discrimination Testing and RepairLili Quan, Tianlin Li, Xiaofei Xie, Zhenpeng Chen 等ICSE 2025 · 被引用 2 次
- 'I Know You Are Discriminatory!': Automated Substantiating for Individual Fairness Auditing of AI SystemsYuanhao Liu, Qi Cao, Huawei Shen, Kaike Zhang 等CSCW 2025
- Constructing a Fair Classifier with Generated Fair DataTaeuk Jang, Feng Zheng, Xiaoqian WangAAAI 2021 · 被引用 44 次
