Fair Bayesian Data Selection via Generalized Discrepancy Measures
Yixuan Zhang, Jiabin Luo, Zhenggang Wang, Feng Zhou, Quyu Kong
Abstract
Fairness concerns are increasingly critical as machine learning models are deployed in high-stakes applications. While existing fairness-aware methods typically intervene at the model level, they often suffer from high computational costs, limited scalability, and poor generalization. To address these challenges, we propose a Bayesian data selection framework that ensures fairness by aligning group-specific posterior distributions of model parameters and sample weights with a shared central distribution. Our framework supports flexible alignment via various distributional discrepancy measures, including Wasserstein distance, maximum mean discrepancy, and f -divergence, allowing geometry-aware control without imposing explicit fairness constraints. This data-centric approach mitigates group-specific biases in training data and improves fairness in downstream tasks, with theoretical guarantees. Experiments on benchmark datasets show that our method consistently outperforms existing data selection and model-based fairness methods in both fairness and accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- FairBatch: Batch Selection for Model FairnessYuji Roh, Kangwook Lee, Steven Euijong Whang, Changho SuhICLR 2021 · 156 citations
- Confidence Scores Make Instance-dependent Label-noise Learning PossibleAntonin Berthon, Bo Han, Gang Niu, Tongliang Liu et al.ICML 2021 · 126 citations
- A General Approach to Fairness with Optimal TransportSilvia Chiappa, Ray Jiang, Tom Stepleton, Aldo Pacchiano et al.AAAI 2020 · 94 citations
- Learning the Stein Discrepancy for Training and Evaluating Energy-Based Models without SamplingWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICML 2020 · 93 citations
- FR-Train: A Mutual Information-Based Approach to Fair and Robust TrainingYuji Roh, Kangwook Lee, Steven Whang, Changho SuhICML 2020 · 90 citations
Related papers
- A Fair Bayesian Inference through Matched Gibbs PosteriorJihu Lee, Kunwoong Kim, Sehyun Park, Insung Kong et al.ICLR 2026
- Fairness via Independence: A General Regularization Framework for Machine LearningYezi Liu, Hanning Chen, Wenjun Huang, Yang Ni et al.ICLR 2026
- Fair and Optimal Classification via Post-ProcessingRuicheng Xian, Lang Yin, Han ZhaoICML 2023 · 57 citations
- Test-Time Debiasing with Probabilistic Prompts via Wasserstein Distance in Vision-Language ModelsChengye Wang, Yuyuan Li, XiaoHua Feng, Xiaolin Zheng et al.ICML 2026
- Can I Trust My Fairness Metric? Assessing Fairness with Unlabeled Data and Bayesian InferenceDisi Ji, Padhraic Smyth, Mark SteyversNeurIPS 2020 · 57 citations
