Experimental Analysis of Multi-Step Pipelines for Fair Classifications - More than the Sum of Their Parts?
Nico Lässig, Melanie Herschel
Abstract
The problem of biased machine learning predictions has led to many alternative approaches to mitigate the problem. They are typically studied and evaluated by focusing on the input data, the trained model, and the performance of the model predictions. We take a broader perspective, considering approaches in the context of a multi-step pipeline. We study fair classification in a pipeline comprising multiple data preparation steps, parameter optimization, and three types of approaches (pre-, in-, and post-processing) designed to reduce bias that may be applied consecutively. This pipeline leads to a trained model to be evaluated in terms of quality (e.g., accuracy) and fairness. We experimentally evaluate the effect differently combined implementations of the pipeline components have on the performance of more than 40 fairness-inducing algorithms. Key findings made possible by this pipeline perspective include: (1) Choosing a bias reducing algorithm greatly simplifies when implementing suited data preparation or parameter optimization, as the difference in performance between methods shrinks, making almost any choice a good one. (2) Several component or pipeline implementations often assumed to have positive or negative effects on performance prove to have little or even contrary effects to the expectations. (3) While many approaches have been published for fair classification in the last decade and shown to improve on previous solutions in specific settings, our broad analysis reveals a stagnating performance trend. Our analysis shows that synergetic effects between pipeline components need to be carefully taken into account for further research on fair end-to-end data processing. It further raises the more fundamental question of how the study of the problem evolves, both in terms of proposed solutions and benchmarking.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get e030caf8-ad60-4b44-951a-3383fc82aaf4Cited by top-tier papers1
Ask how each one uses itRelated papers
- FRAPPÉ: A Group Fairness Framework for Post-Processing EverythingAlexandru Tifrea, Preethi Lahoti, Ben Packer, Yoni Halpern et al.ICML 2024 · 15 citations
- Fair preprocessing: towards understanding compositional fairness of data transformers in machine learning pipelineSumon Biswas, Hridesh RajanFSE 2021 · 101 citations
- Demystifying the Optimal Fair Classifier in Multi-Class ClassificationLi Zhang, Yuyuan Li, XiaoHua Feng, Jiaming Zhang et al.ICML 2026
- Through the Data Management Lens: Experimental Analysis and Evaluation of Fair ClassificationMaliha Tashfia Islam, Anna Fariha, Alexandra Meliou, Babak SalimiSIGMOD 2022 · 29 citations
- Beyond Adult and COMPAS: Fair Multi-Class Prediction via Information ProjectionWael Alghamdi, Hsiang Hsu, Haewon Jeong, Hao Wang et al.NeurIPS 2022 · 57 citations
