Navigating Towards Fairness with Data Selection
Yixuan Zhang, Zhidong Li, Yang Wang, Fang Chen, Xuhui Fan, Feng Zhou
Abstract
Machine learning algorithms often struggle to eliminate inherent data biases, particularly those arising from unreliable labels, which poses a significant challenge in ensuring fairness. Existing fairness techniques that address label bias typically involve modifying models and intervening in the training process, but these lack flexibility for large-scale datasets. To address this limitation, we introduce a data selection method designed to efficiently and flexibly mitigate label bias, tailored to more practical needs. Our approach utilizes a zero-shot predictor as a proxy model that simulates training on a clean holdout set. This strategy, supported by peer predictions, ensures the fairness of the proxy model and eliminates the need for an additional holdout set, which is a common requirement in previous methods. Without altering the classifier's architecture, our modality-agnostic method effectively selects appropriate training data and has proven efficient and effective in handling label bias and improving fairness across diverse datasets in experimental evaluations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
- GLISTER: Generalization based Data Subset Selection for Efficient and Robust LearningKrishnaTeja Killamsetty, Durga Sivasubramanian, Ganesh Ramakrishnan, Rishabh K. IyerAAAI 2021 · 300 citations
- Peer Loss Functions: Learning from Noisy Labels without Knowing Noise RatesYang Liu, Hongyi GuoICML 2020 · 280 citations
- Prioritized Training on Points that are Learnable, Worth Learning, and not yet LearntSören Mindermann, Jan Markus Brauner, Muhammed Razzak, Mrinank Sharma et al.ICML 2022 · 237 citations
- Diagnosing failures of fairness transfer across distribution shift in real-world medical settingsJessica Schrouff, Natalie Harris, Sanmi Koyejo, Ibrahim M. Alabdulmohsin et al.NeurIPS 2022 · 84 citations
Related papers
- Towards Accelerated Model Training via Bayesian Data SelectionZhijie Deng, Peng Cui, Jun ZhuNeurIPS 2023 · 15 citations
- Mitigating Label Bias in Machine Learning: Fairness through Confident LearningYixuan Zhang, Boyu Li, Zenan Ling, Feng ZhouAAAI 2024 · 8 citations
- Constructing a Fair Classifier with Generated Fair DataTaeuk Jang, Feng Zheng, Xiaoqian WangAAAI 2021 · 44 citations
- Are Your Fairness Metrics Accurate? A Semi-Supervised Approach to Improving Fairness Estimates Under Sample Selection BiasM. Clara De Paolis Kaluza, Thulasi Tholeti, Yile Chen, Ricardo Baeza-Yates et al.KDD 2025 · 1 citation
- Fair Infinitesimal Jackknife: Mitigating the Influence of Biased Training Data Points Without RefittingPrasanna Sattigeri, Soumya Ghosh, Inkit Padhi, Pierre L. Dognin et al.NeurIPS 2022 · 36 citations
