Fair preprocessing: towards understanding compositional fairness of data transformers in machine learning pipeline
Sumon Biswas, Hridesh Rajan
Abstract
In recent years, many incidents have been reported where machine learning models exhibited discrimination among people based on race, sex, age, etc. Research has been conducted to measure and mitigate unfairness in machine learning models. For a machine learning task, it is a common practice to build a pipeline that includes an ordered set of data preprocessing stages followed by a classifier. However, most of the research on fairness has considered a single classifier based prediction task. What are the fairness impacts of the preprocessing stages in machine learning pipeline? Furthermore, studies showed that often the root cause of unfairness is ingrained in the data itself, rather than the model. But no research has been conducted to measure the unfairness caused by a specific transformation made in the data preprocessing stage. In this paper, we introduced the causal method of fairness to reason about the fairness impact of data preprocessing stages in ML pipeline. We leveraged existing metrics to define the fairness measures of the stages. Then we conducted a detailed fairness evaluation of the preprocessing stages in 37 pipelines collected from three different sources. Our results show that certain data transformers are causing the model to exhibit unfairness. We identified a number of fairness patterns in several categories of data transformers. Finally, we showed how the local fairness of a preprocessing stage composes in the global fairness of the pipeline. We used the fairness composition to choose appropriate downstream transformer that mitigates unfairness in the machine learning pipeline.
• Software and its engineering → Software creation and management; • Computing methodologies → Machine learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3d04c438-5d2e-4ffa-87c9-120168546397Cited by top-tier papers25
- Fairea: a model behaviour mutation approach to benchmarking bias mitigation methodsMax Hort, Jie M. Zhang, Federica Sarro, Mark HarmanFSE 2021 · 75 citations
- MAAT: a novel ensemble approach to addressing fairness and performance bugs for machine learning softwareZhenpeng Chen, Jie M. Zhang, Federica Sarro, Mark HarmanFSE 2022 · 65 citations
- The Art and Practice of Data Science Pipelines: A Comprehensive Study of Data Science Pipelines In Theory, In-The-Small, and In-The-LargeSumon Biswas, Mohammad Wardat, Hridesh RajanICSE 2022 · 64 citations
- NeuronFair: Interpretable White-Box Fairness Testing through Biased Neuron IdentificationHaibin Zheng, Zhiqing Chen, Tianyu Du, Xuhong Zhang et al.ICSE 2022 · 58 citations
- Fairness Improvement with Multiple Protected Attributes: How Far Are We?Zhenpeng Chen, Jie M. Zhang, Federica Sarro, Mark HarmanICSE 2024 · 33 citations
Builds on6
- Repairing deep neural networks: fix patterns and challengesMd Johirul Islam, Rangeet Pan, Giang Nguyen, Hridesh RajanICSE 2020 · 102 citations
- Do the machine learning models on a crowd sourced platform exhibit bias? an empirical study on model fairnessSumon Biswas, Hridesh RajanFSE 2020 · 96 citations
- DeepLocalize: Fault Localization for Deep Neural NetworksMohammad Wardat, Wei Le, Hridesh RajanICSE 2021 · 93 citations
- Fairea: a model behaviour mutation approach to benchmarking bias mitigation methodsMax Hort, Jie M. Zhang, Federica Sarro, Mark HarmanFSE 2021 · 75 citations
- "Ignorance and Prejudice" in Software FairnessJie M. Zhang, Mark HarmanICSE 2021 · 69 citations
Related papers
- Experimental Analysis of Multi-Step Pipelines for Fair Classifications - More than the Sum of Their Parts?Nico Lässig, Melanie HerschelICDE 2025
- Causality-Aided Trade-Off Analysis for Machine Learning FairnessZhenlan Ji, Pingchuan Ma, Shuai Wang, Yanhui LiASE 2023 · 6 citations
- Fairway: a way to build fair ML softwareJoymallya Chakraborty, Suvodeep Majumder, Zhe Yu, Tim MenziesFSE 2020 · 131 citations
- Adaptive fairness improvement based on causality analysisMengdi Zhang, Jun SunFSE 2022 · 35 citations
- Capturing and querying fine-grained provenance of preprocessing pipelines in data scienceAdriane Chapman, Paolo Missier, Giulia Simonelli, Riccardo TorloneVLDB 2021 · 39 citations
