Towards Understanding Fairness and its Composition in Ensemble Machine Learning
Usman Gohar, Sumon Biswas, Hridesh Rajan
摘要
Machine Learning (ML) software has been widely adopted in modern society, with reported fairness implications for minority groups based on race, sex, age, etc. Many recent works have proposed methods to measure and mitigate algorithmic bias in ML models. The existing approaches focus on single classifier-based ML models. However, real-world ML models are often composed of multiple independent or dependent learners in an ensemble (e.g., Random Forest), where the fairness composes in a non-trivial way. How does fairness compose in ensembles? What are the fairness impacts of the learners on the ultimate fairness of the ensemble? Can fair learners result in an unfair ensemble? Furthermore, studies have shown that hyperparameters influence the fairness of ML models. Ensemble hyperparameters are more complex since they affect how learners are combined in different categories of ensembles. Understanding the impact of ensemble hyperparameters on fairness will help programmers design fair ensembles. Today, we do not understand these fully for different ensemble algorithms. In this paper, we comprehensively study popular real-world ensembles: Bagging, Boosting, Stacking, and Voting. We have developed a benchmark of 168 ensemble models collected from Kaggle on four popular fairness datasets. We use existing fairness metrics to understand the composition of fairness. Our results show that ensembles can be designed to be fairer without using mitigation techniques. We also identify the interplay between fairness composition and data characteristics to guide fair ensemble design. Finally, our benchmark can be leveraged for further research on fair ensembles. To the best of our knowledge, this is one of the first and largest studies on fairness composition in ensembles yet presented in the literature.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Fairness Improvement with Multiple Protected Attributes: How Far Are We?Zhenpeng Chen, Jie M. Zhang, Federica Sarro, Mark HarmanICSE 2024 · 被引用 33 次
- Fix Fairness, Don't Ruin Accuracy: Performance Aware Fairness Repair using AutoMLGiang Nguyen, Sumon Biswas, Hridesh RajanFSE 2023 · 被引用 15 次
- On Fairness of Task Arithmetic: The Role of Task VectorsLaura Gomezjurado Gonzalez, Hiroki Naganuma, Kotaro Yoshida, Takafumi Horie 等ICLR 2026 · 被引用 3 次
- Inferring Data Preconditions from Deep Learning Models for Trustworthy Prediction in DeploymentShibbir Ahmed, Hongyang Gao, Hridesh RajanICSE 2024 · 被引用 3 次
- Dissecting Global Search: A Simple Yet Effective Method to Boost Individual Discrimination Testing and RepairLili Quan, Tianlin Li, Xiaofei Xie, Zhenpeng Chen 等ICSE 2025 · 被引用 2 次
它引用的顶会 Paper11
- Bias in machine learning software: why? how? what to do?Joymallya Chakraborty, Suvodeep Majumder, Tim MenziesFSE 2021 · 被引用 186 次
- Fairway: a way to build fair ML softwareJoymallya Chakraborty, Suvodeep Majumder, Zhe Yu, Tim MenziesFSE 2020 · 被引用 131 次
- Fair preprocessing: towards understanding compositional fairness of data transformers in machine learning pipelineSumon Biswas, Hridesh RajanFSE 2021 · 被引用 101 次
- Do the machine learning models on a crowd sourced platform exhibit bias? an empirical study on model fairnessSumon Biswas, Hridesh RajanFSE 2020 · 被引用 96 次
- "Ignorance and Prejudice" in Software FairnessJie M. Zhang, Mark HarmanICSE 2021 · 被引用 69 次
相关 Paper
- MAAT: a novel ensemble approach to addressing fairness and performance bugs for machine learning softwareZhenpeng Chen, Jie M. Zhang, Federica Sarro, Mark HarmanFSE 2022 · 被引用 65 次
- Individual Arbitrariness and Group FairnessCarol Xuan Long, Hsiang Hsu, Wael Alghamdi, Flávio P. CalmonNeurIPS 2023 · 被引用 16 次
- Diversity Drives Fairness: Ensemble of Higher Order Mutants for Intersectional Fairness of Machine Learning SoftwareZhenpeng Chen, Xinyue Li, Jie M. Zhang, Federica Sarro 等ICSE 2025 · 被引用 2 次
- Fairness-aware Configuration of Machine Learning LibrariesSaeid Tizpaz-Niari, Ashish Kumar, Gang Tan, Ashutosh TrivediICSE 2022 · 被引用 44 次
- StackGenVis: Alignment of Data, Algorithms, and Models for Stacking Ensemble Learning Using Performance MetricsAngelos Chatzimparmpas, Rafael Messias Martins, Kostiantyn Kucher, Andreas KerrenIEEE VIS 2020 · 被引用 153 次
