Are My Deep Learning Systems Fair? An Empirical Study of Fixed-Seed Training
Shangshu Qian, Hung Viet Pham, Thibaud Lutellier, Zeou Hu, Jungwon Kim, Lin Tan, Yaoliang Yu, Jiahao Chen, Sameena Shah
Abstract
Deep learning (DL) systems have been gaining popularity in critical tasks such as credit evaluation and crime prediction. Such systems demand fairness. Recent work shows that DL software implementations introduce variance: identical DL training runs (i.e., identical network, data, configuration, software, and hardware) with a fixed seed produce different models. Such variance could make DL models and networks violate fairness compliance laws, resulting in negative social impact. In this paper, we conduct the first empirical study to quantify the impact of software implementation on the fairness and its variance of DL systems. Our study of 22 mitigation techniques and five baselines reveals up to 12.6% fairness variance across identical training runs with identical seeds. In addition, most debiasing algorithms have a negative impact on the model such as reducing model accuracy, increasing fairness variance, or increasing accuracy variance. Our literature survey shows that while fairness is gaining popularity in artificial intelligence (AI) related conferences, only 34.4% of the papers use multiple identical training runs to evaluate their approach, raising concerns about their results' validity. We call for better fairness evaluation and testing protocols to improve fairness and fairness variance of DL systems as well as DL research validity and reproducibility at large. * Jiahao Chen has since moved to Parity Technologies. 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e9fab9d2-6c1b-4403-aa58-a0946e9f9510Cited by top-tier papers6
- Trivial or Impossible --- dichotomous data difficulty masks model differences (on ImageNet and beyond)Kristof Meding, Luca M. Schulze Buschoff, Robert Geirhos, Felix A. WichmannICLR 2022 · 47 citations
- DisGUIDE: Disagreement-Guided Data-Free Model ExtractionJonathan Rosenthal, Eric Enouen, Hung Viet Pham, Lin TanAAAI 2023 · 31 citations
- Input-agnostic Certified Group Fairness via Gaussian Parameter SmoothingJiayin Jin, Zeru Zhang, Yang Zhou, Lingfei WuICML 2022 · 18 citations
- EasyScale: Elastic Training with Consistent Accuracy and Improved Utilization on GPUsMingzhen Li, Wencong Xiao, Hailong Yang, Biao Sun et al.SC 2023 · 16 citations
- Mixtera: A Data Plane for Foundation Model TrainingMaximilian Böther, Xiaozhe Yao, Tolga Kerimoglu, Dan Graur et al.SIGMOD 2026 · 5 citations
Builds on25
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang et al.ICCV 2019 · 469 citations
- Learning from Failure: De-biasing Classifier from Biased ClassifierJun Hyun Nam, Hyuntak Cha, Sungsoo Ahn, Jaeho Lee et al.NeurIPS 2020 · 428 citations
- On Measuring and Mitigating Biased Inferences of Word EmbeddingsSunipa Dev, Tao Li, Jeff M. Phillips, Vivek SrikumarAAAI 2020 · 195 citations
- Bias in machine learning software: why? how? what to do?Joymallya Chakraborty, Suvodeep Majumder, Tim MenziesFSE 2021 · 186 citations
Related papers
- Problems and Opportunities in Training Deep Learning Software Systems: An Analysis of VarianceHung Viet Pham, Shangshu Qian, Jiannan Wang, Thibaud Lutellier et al.ASE 2020 · 91 citations
- A Large-Scale Empirical Study on Improving the Fairness of Image Classification ModelsJunjie Yang, Jiajun Jiang, Zeyu Sun, Junjie ChenISSTA 2024 · 4 citations
- Do the machine learning models on a crowd sourced platform exhibit bias? an empirical study on model fairnessSumon Biswas, Hridesh RajanFSE 2020 · 96 citations
- Fairea: a model behaviour mutation approach to benchmarking bias mitigation methodsMax Hort, Jie M. Zhang, Federica Sarro, Mark HarmanFSE 2021 · 75 citations
- Towards Understanding Fairness and its Composition in Ensemble Machine LearningUsman Gohar, Sumon Biswas, Hridesh RajanICSE 2023 · 30 citations
