From Biased Selective Labels to Pseudo-Labels: An Expectation-Maximization Framework for Learning from Biased Decisions
Trenton Chang, Jenna Wiens
摘要
Selective labels occur when label observations are subject to a decision-making process; e.g., diagnoses that depend on the administration of laboratory tests. We study a clinically-inspired selective label problem called disparate censorship, where labeling biases vary across subgroups and unlabeled individuals are imputed as "negative" (i.e., no diagnostic test = no illness). Machine learning models naïvely trained on such labels could amplify labeling bias. Inspired by causal models of selective labels, we propose Disparate Censorship Expectation-Maximization (DCEM), an algorithm for learning in the presence of disparate censorship. We theoretically analyze how DCEM mitigates the effects of disparate censorship on model performance. We validate DCEM on synthetic data, showing that it improves bias mitigation (area between ROC curves) without sacrificing discriminative performance (AUC) compared to baselines. We achieve similar results in a sepsis classification task using clinical data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper16
- In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised LearningMamshad Nayeem Rizve, Kevin Duarte, Yogesh S. Rawat, Mubarak ShahICLR 2021 · 被引用 630 次
- SELF: Learning to Filter Noisy Labels with Self-EnsemblingDuc Tam Nguyen, Chaithanya Kumar Mummadi, Thi-Phuong-Nhung Ngo, Thi Hoai Phuong Nguyen 等ICLR 2020 · 被引用 354 次
- Peer Loss Functions: Learning from Noisy Labels without Knowing Noise RatesYang Liu, Hongyi GuoICML 2020 · 被引用 280 次
- Learning with Feature-Dependent Label Noise: A Progressive ApproachYikai Zhang, Songzhu Zheng, Pengxiang Wu, Mayank Goswami 等ICLR 2021 · 被引用 184 次
- Generalized Jensen-Shannon Divergence Loss for Learning with Noisy LabelsErik Englesson, Hossein AzizpourNeurIPS 2021 · 被引用 170 次
相关 Paper
- Interventional Multi-Instance Learning with Deconfounded Instance-Level PredictionTiancheng Lin, Hongteng Xu, Canqian Yang, Yi XuAAAI 2022 · 被引用 34 次
- Are labels informative in semi-supervised learning? Estimating and leveraging the missing-data mechanismAude Sportisse, Hugo Schmutz, Olivier Humbert, Charles Bouveyron 等ICML 2023 · 被引用 10 次
- Fair Selective Classification Via SufficiencyJoshua K. Lee, Yuheng Bu, Deepta Rajan, Prasanna Sattigeri 等ICML 2021 · 被引用 33 次
- Are Your Fairness Metrics Accurate? A Semi-Supervised Approach to Improving Fairness Estimates Under Sample Selection BiasM. Clara De Paolis Kaluza, Thulasi Tholeti, Yile Chen, Ricardo Baeza-Yates 等KDD 2025 · 被引用 1 次
- The Rich Get Richer: Disparate Impact of Semi-Supervised LearningZhaowei Zhu, Tianyi Luo, Yang LiuICLR 2022 · 被引用 44 次
