FairWASP: Fast and Optimal Fair Wasserstein Pre-processing
Zikai Xiong, Niccolò Dalmasso, Alan Mishler, Vamsi K. Potluru, Tucker Balch, Manuela Veloso
Abstract
Recent years have seen a surge of machine learning approaches aimed at reducing disparities in model outputs across different subgroups. In many settings, training data may be used in multiple downstream applications by different users, which means it may be most effective to intervene on the training data itself. In this work, we present FairWASP, a novel pre-processing approach designed to reduce disparities in classification datasets without modifying the original data. FairWASP returns sample-level weights such that the reweighted dataset minimizes the Wasserstein distance to the original dataset while satisfying (an empirical version of) demographic parity, a popular fairness criterion. We show theoretically that integer weights are optimal, which means our method can be equivalently understood as duplicating or eliminating samples. FairWASP can therefore be used to construct datasets which can be fed into any classification method, not just methods which accept sample weights. Our work is based on reformulating the pre-processing task as a large-scale mixed-integer program (MIP), for which we propose a highly efficient algorithm based on the cutting plane method. Experiments demonstrate that our proposed optimization algorithm significantly outperforms state-of-the-art commercial solvers in solving both the MIP and its linear program relaxation. Further experiments highlight the competitive performance of FairWASP in reducing disparities while preserving accuracy in downstream classification settings. * Corresponding Author straints or regularizers during the model training process itself, and (iii) post-processing methods alter the outputs of previously trained models. See Hort et al. ( 2022 ) for a recent review of methods across all three categories. Among these three, no one category of methods clearly dominates the others in terms of performance. Preprocessing methods are useful when the person who generates or maintains a dataset is not the same as the person who will be using it to train a model (Feldman et al. 2015) , or when a dataset may be used to train multiple models. These methods typically require no knowledge of downstream models, so they are in principle compatible with any subsequent machine learning procedure. Many pre-processing methods operate by changing the feature values or labels of the training data (Calders, Kami-
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12b2f253-ebab-48ac-a994-4c70271e7db1Cited by top-tier papers6
- Fair Wasserstein CoresetsZikai Xiong, Niccolò Dalmasso, Shubham Sharma, Freddy Lécué et al.NeurIPS 2024 · 10 citations
- Understanding Accuracy-Fairness Trade-offs in Re-ranking through Elasticity in EconomicsChen Xu, Jujia Zhao, Wenjie Wang, Liang Pang et al.SIGIR 2025 · 4 citations
- On Optimal Steering to Achieve Exact FairnessMohit Sharma, Amit Deshpande, Chiranjib Bhattacharyya, Rajiv Ratn ShahNeurIPS 2025 · 2 citations
- Fair Data Pre-Processing with Imperfect Attribute SpaceYing Zheng, Yangfan Jiang, Kian-Lee TanSIGMOD 2026
- Bridging Jensen Gap for Max-Min Group Fairness Optimization in RecommendationChen Xu, Yuxin Li, Wenjie Wang, Liang Pang et al.ICLR 2025
Builds on6
- Bias in machine learning software: why? how? what to do?Joymallya Chakraborty, Suvodeep Majumder, Tim MenziesFSE 2021 · 186 citations
- Practical Large-Scale Linear Programming using Primal-Dual Hybrid GradientDavid L. Applegate, Mateo Díaz, Oliver Hinder, Haihao Lu et al.NeurIPS 2021 · 165 citations
- Achieving Fairness at No Utility Cost via Data Reweighing with InfluencePeizhao Li, Hongfu LiuICML 2022 · 57 citations
- An improved cutting plane method for convex optimization, convex-concave games, and its applicationsHaotian Jiang, Yin Tat Lee, Zhao Song, Sam Chiu-wai WongSTOC 2020 · 54 citations
- Fairness with Adaptive WeightsJunyi Chai, Xiaoqian WangICML 2022 · 47 citations
Related papers
- Fair Infinitesimal Jackknife: Mitigating the Influence of Biased Training Data Points Without RefittingPrasanna Sattigeri, Soumya Ghosh, Inkit Padhi, Pierre L. Dognin et al.NeurIPS 2022 · 36 citations
- Fair and Optimal Classification via Post-ProcessingRuicheng Xian, Lang Yin, Han ZhaoICML 2023 · 57 citations
- FairBatch: Batch Selection for Model FairnessYuji Roh, Kangwook Lee, Steven Euijong Whang, Changho SuhICLR 2021 · 156 citations
- Exploiting MMD and Sinkhorn Divergences for Fair and Transferable Representation LearningLuca Oneto, Michele Donini, Giulia Luise, Carlo Ciliberto et al.NeurIPS 2020 · 56 citations
- Experimental Analysis of Multi-Step Pipelines for Fair Classifications - More than the Sum of Their Parts?Nico Lässig, Melanie HerschelICDE 2025
