The Rich Get Richer: Disparate Impact of Semi-Supervised Learning
Zhaowei Zhu, Tianyi Luo, Yang Liu
摘要
Semi-supervised learning (SSL) has demonstrated its potential to improve the model accuracy for a variety of learning tasks when the high-quality supervised data is severely limited. Although it is often established that the average accuracy for the entire population of data is improved, it is unclear how SSL fares with different sub-populations. Understanding the above question has substantial fairness implications when different sub-populations are defined by the demographic groups that we aim to treat fairly. In this paper, we reveal the disparate impacts of deploying SSL: the sub-population who has a higher baseline accuracy without using SSL (the "rich" one) tends to benefit more from SSL; while the sub-population who suffers from a low baseline accuracy (the "poor" one) might even observe a performance drop after adding the SSL module. We theoretically and empirically establish the above observation for a broad family of SSL algorithms, which either explicitly or implicitly use an auxiliary "pseudo-label". Experiments on a set of image and text classification tasks confirm our claims. We introduce a new metric, Benefit Ratio, and promote the evaluation of the fairness of SSL (Equalized Benefit Ratio). We further discuss how the disparate impact can be mitigated. We hope our paper will alarm the potential pitfall of using SSL and encourage a multifaceted evaluation of future SSL algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Detecting Corrupted Labels Without Training a Model to PredictZhaowei Zhu, Zihao Dong, Yang LiuICML 2022 · 被引用 84 次
- Data Feedback Loops: Model-driven Amplification of Dataset BiasesRohan Taori, Tatsunori HashimotoICML 2023 · 被引用 67 次
- Pruning has a disparate impact on model accuracyCuong Tran, Ferdinando Fioretto, Jung-Eun Kim, Rakshit NaiduNeurIPS 2022 · 被引用 64 次
- Transferring Fairness under Distribution Shifts via Fair Consistency RegularizationBang An, Zora Che, Mucong Ding, Furong HuangNeurIPS 2022 · 被引用 44 次
- Beyond Images: Label Noise Transition Matrix Estimation for Tasks with Lower-Quality FeaturesZhaowei Zhu, Jialu Wang, Yang LiuICML 2022 · 被引用 43 次
它引用的顶会 Paper24
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- Early-Learning Regularization Prevents Memorization of Noisy LabelsSheng Liu, Jonathan Niles-Weed, Narges Razavian, Carlos Fernandez-GrandaNeurIPS 2020 · 被引用 798 次
- Fairness without Demographics through Adversarially Reweighted LearningPreethi Lahoti, Alex Beutel, Jilin Chen, Kang Lee 等NeurIPS 2020 · 被引用 406 次
- MixText: Linguistically-Informed Interpolation of Hidden Space for Semi-Supervised Text ClassificationJiaao Chen, Zichao Yang, Diyi YangACL 2020 · 被引用 340 次
- Learning with Noisy Labels Revisited: A Study Using Real-World Human AnnotationsJiaheng Wei, Zhaowei Zhu, Hao Cheng, Tongliang Liu 等ICLR 2022 · 被引用 338 次
相关 Paper
- Distribution Aligning Refinery of Pseudo-label for Imbalanced Semi-supervised LearningJaehyung Kim, Youngbum Hur, Sejun Park, Eunho Yang 等NeurIPS 2020 · 被引用 209 次
- CDMAD: Class-Distribution-Mismatch-Aware Debiasing for Class-Imbalanced Semi-Supervised LearningHyuck Lee, Heeyoung KimCVPR 2024
- Are Your Fairness Metrics Accurate? A Semi-Supervised Approach to Improving Fairness Estimates Under Sample Selection BiasM. Clara De Paolis Kaluza, Thulasi Tholeti, Yile Chen, Ricardo Baeza-Yates 等KDD 2025 · 被引用 1 次
- Class-Imbalanced Semi-Supervised Learning with Adaptive ThresholdingLan-Zhe Guo, Yufeng LiICML 2022 · 被引用 148 次
- ABC: Auxiliary Balanced Classifier for Class-imbalanced Semi-supervised LearningHyuck Lee, Seungjae Shin, Heeyoung KimNeurIPS 2021 · 被引用 131 次
