Fair Wasserstein Coresets
Zikai Xiong, Niccolò Dalmasso, Shubham Sharma, Freddy Lécué, Daniele Magazzeni, Vamsi K. Potluru, Tucker Balch, Manuela Veloso
摘要
Data distillation and coresets have emerged as popular approaches to generate a smaller representative set of samples for downstream learning tasks to handle large-scale datasets. At the same time, machine learning is being increasingly applied to decision-making processes at a societal level, making it imperative for modelers to address inherent biases towards subgroups present in the data. While current approaches focus on creating fair synthetic representative samples by optimizing local properties relative to the original samples, their impact on downstream learning processes has yet to be explored. In this work, we present fair Wasserstein coresets (FWC), a novel coreset approach which generates fair synthetic representative samples along with sample-level weights to be used in downstream learning tasks. FWC uses an efficient majority minimization algorithm to minimize the Wasserstein distance between the original dataset and the weighted synthetic samples while enforcing demographic parity. We show that an unconstrained version of FWC is equivalent to Lloyd's algorithm for k-medians and k-means clustering. Experiments conducted on both synthetic and real datasets show that FWC: (i) achieves a competitive fairness-utility tradeoff in downstream models compared to existing approaches, (ii) improves downstream fairness when added to the existing training data and (iii) can be used to reduce biases in predictions from large language models (GPT-3.5 and GPT-4).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Fair Dataset Distillation via Cross-Group Barycenter AlignmentMohammad Hossein Moslemi, Nima Hosseini Dashtbayaz, Zhimin Mei, Bissan Ghaddar 等ICML 2026
- Optimal Transport under Group Fairness ConstraintsLinus Bleistein, Mathieu Dagréou, Francisco Andrade, Thomas Boudou 等ICML 2026
它引用的顶会 Paper12
- Dataset Condensation with Gradient MatchingBo Zhao, Konda Reddy Mopuri, Hakan BilenICLR 2021 · 被引用 684 次
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 被引用 494 次
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman 等ICLR 2020 · 被引用 462 次
- Coresets via Bilevel Optimization for Continual Learning and StreamingZalán Borsos, Mojmir Mutny, Andreas KrauseNeurIPS 2020 · 被引用 320 次
- Explainable k-Means and k-Medians ClusteringMichal Moshkovitz, Sanjoy Dasgupta, Cyrus Rashtchian, Nave FrostICML 2020 · 被引用 184 次
相关 Paper
- On Optimal Steering to Achieve Exact FairnessMohit Sharma, Amit Deshpande, Chiranjib Bhattacharyya, Rajiv Ratn ShahNeurIPS 2025 · 被引用 2 次
- FairWASP: Fast and Optimal Fair Wasserstein Pre-processingZikai Xiong, Niccolò Dalmasso, Alan Mishler, Vamsi K. Potluru 等AAAI 2024 · 被引用 7 次
- Finding Wasserstein Ball Center: Efficient Algorithm and The Applications in FairnessYuntao Wang, Yuxuan Li, Qingyuan Yang, Hu DingICML 2025
- Conditional Learning of Fair RepresentationsHan Zhao, Amanda Coston, Tameem Adel, Geoffrey J. GordonICLR 2020 · 被引用 127 次
- Exploiting MMD and Sinkhorn Divergences for Fair and Transferable Representation LearningLuca Oneto, Michele Donini, Giulia Luise, Carlo Ciliberto 等NeurIPS 2020 · 被引用 56 次
