"Ignorance and Prejudice" in Software Fairness
Jie M. Zhang, Mark Harman
Abstract
Machine learning software can be unfair when making human-related decisions, having prejudices over certain groups of people. Existing work primarily focuses on proposing fairness metrics and presenting fairness improvement approaches. It remains unclear how key aspect of any machine learning system, such as feature set and training data, affect fairness. This paper presents results from a comprehensive study that addresses this problem. We find that enlarging the feature set plays a significant role in fairness (with an average effect rate of 38%). Importantly, and contrary to widely-held beliefs that greater fairness often corresponds to lower accuracy, our findings reveal that an enlarged feature set has both higher accuracy and fairness. Perhaps also surprisingly, we find that a larger training data does not help to improve fairness. Our results suggest a larger training data set has more unfairness than a smaller one when feature sets are insufficient; an important cautionary finding for practising software engineers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e89ffdc4-2018-446c-af80-bde7a7678452Cited by top-tier papers21
- Fair preprocessing: towards understanding compositional fairness of data transformers in machine learning pipelineSumon Biswas, Hridesh RajanFSE 2021 · 101 citations
- Fairea: a model behaviour mutation approach to benchmarking bias mitigation methodsMax Hort, Jie M. Zhang, Federica Sarro, Mark HarmanFSE 2021 · 75 citations
- MAAT: a novel ensemble approach to addressing fairness and performance bugs for machine learning softwareZhenpeng Chen, Jie M. Zhang, Federica Sarro, Mark HarmanFSE 2022 · 65 citations
- NeuronFair: Interpretable White-Box Fairness Testing through Biased Neuron IdentificationHaibin Zheng, Zhiqing Chen, Tianyu Du, Xuhong Zhang et al.ICSE 2022 · 58 citations
- Information-Theoretic Testing and Debugging of Fairness Defects in Deep Neural NetworksVerya Monjezi, Ashutosh Trivedi, Gang Tan, Saeid Tizpaz-NiariICSE 2023 · 47 citations
Builds on3
- Fairway: a way to build fair ML softwareJoymallya Chakraborty, Suvodeep Majumder, Zhe Yu, Tim MenziesFSE 2020 · 131 citations
- Automatic testing and improvement of machine translationZeyu Sun, Jie M. Zhang, Mark Harman, Mike Papadakis et al.ICSE 2020 · 111 citations
- Do the machine learning models on a crowd sourced platform exhibit bias? an empirical study on model fairnessSumon Biswas, Hridesh RajanFSE 2020 · 96 citations
Related papers
- Fairness Improvement with Multiple Protected Attributes: How Far Are We?Zhenpeng Chen, Jie M. Zhang, Federica Sarro, Mark HarmanICSE 2024 · 33 citations
- Towards Understanding Fairness and its Composition in Ensemble Machine LearningUsman Gohar, Sumon Biswas, Hridesh RajanICSE 2023 · 30 citations
- Training Data Debugging for the Fairness of Machine Learning SoftwareYanhui Li, Linghan Meng, Lin Chen, Li Yu et al.ICSE 2022 · 49 citations
- A Large-Scale Empirical Study on Improving the Fairness of Image Classification ModelsJunjie Yang, Jiajun Jiang, Zeyu Sun, Junjie ChenISSTA 2024 · 4 citations
- Constructing a Fair Classifier with Generated Fair DataTaeuk Jang, Feng Zheng, Xiaoqian WangAAAI 2021 · 44 citations
