Bias in machine learning software: why? how? what to do?
Joymallya Chakraborty, Suvodeep Majumder, Tim Menzies
Abstract
Increasingly, software is making autonomous decisions in case of criminal sentencing, approving credit cards, hiring employees, and so on. Some of these decisions show bias and adversely affect certain social groups (e.g. those defined by sex, race, age, marital status). Many prior works on bias mitigation take the following form: change the data or learners in multiple ways, then see if any of that improves fairness. Perhaps a better approach is to postulate root causes of bias and then applying some resolution strategy.
This paper checks if the root causes of bias are the prior decisions about (a) what data was selected and (b) the labels assigned to those examples. Our Fair-SMOTE algorithm removes biased labels; and rebalances internal distributions so that, based on sensitive attribute, examples are equal in positive and negative classes. On testing, this method was just as effective at reducing bias as prior approaches. Further, models generated via Fair-SMOTE achieve higher performance (measured in terms of recall and F1) than other state-of-the-art fairness improvement algorithms.
To the best of our knowledge, measured in terms of number of analyzed learners and datasets, this study is one of the largest studies on bias mitigation yet presented in the literature.
• Software and its engineering → Software creation and management; • Computing methodologies → Machine learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e742ba91-0458-4f54-9389-c92911df3ef4Cited by top-tier papers26
- MAAT: a novel ensemble approach to addressing fairness and performance bugs for machine learning softwareZhenpeng Chen, Jie M. Zhang, Federica Sarro, Mark HarmanFSE 2022 · 65 citations
- Are My Deep Learning Systems Fair? An Empirical Study of Fixed-Seed TrainingShangshu Qian, Hung Viet Pham, Thibaud Lutellier, Zeou Hu et al.NeurIPS 2021 · 49 citations
- Information-Theoretic Testing and Debugging of Fairness Defects in Deep Neural NetworksVerya Monjezi, Ashutosh Trivedi, Gang Tan, Saeid Tizpaz-NiariICSE 2023 · 47 citations
- Explanation-Guided Fairness Testing through Genetic AlgorithmMing Fan, Wenying Wei, Wuxia Jin, Zijiang Yang et al.ICSE 2022 · 45 citations
- Fairness Improvement with Multiple Protected Attributes: How Far Are We?Zhenpeng Chen, Jie M. Zhang, Federica Sarro, Mark HarmanICSE 2024 · 33 citations
Builds on3
- Fairway: a way to build fair ML softwareJoymallya Chakraborty, Suvodeep Majumder, Zhe Yu, Tim MenziesFSE 2020 · 131 citations
- White-box fairness testing through adversarial samplingPeixin Zhang, Jingyi Wang, Jun Sun, Guoliang Dong et al.ICSE 2020 · 127 citations
- Do the machine learning models on a crowd sourced platform exhibit bias? an empirical study on model fairnessSumon Biswas, Hridesh RajanFSE 2020 · 96 citations
Related papers
- Fairea: a model behaviour mutation approach to benchmarking bias mitigation methodsMax Hort, Jie M. Zhang, Federica Sarro, Mark HarmanFSE 2021 · 75 citations
- Fix Fairness, Don't Ruin Accuracy: Performance Aware Fairness Repair using AutoMLGiang Nguyen, Sumon Biswas, Hridesh RajanFSE 2023 · 15 citations
- Fair Infinitesimal Jackknife: Mitigating the Influence of Biased Training Data Points Without RefittingPrasanna Sattigeri, Soumya Ghosh, Inkit Padhi, Pierre L. Dognin et al.NeurIPS 2022 · 36 citations
- On the Robustness of Fairness Practices: A Causal Framework for Systematic EvaluationVerya Monjezi, Ashish Kumar, Ashutosh Trivedi, Gang Tan et al.ICSE 2026
- Adaptive fairness improvement based on causality analysisMengdi Zhang, Jun SunFSE 2022 · 35 citations
