Mitigating Source Bias for Fairer Weak Supervision
Changho Shin, Sonia Cromp, Dyah Adila, Frederic Sala
摘要
Weak supervision enables efficient development of training sets by reducing the need for ground truth labels. However, the techniques that make weak supervision attractive -- such as integrating any source of signal to estimate unknown labels -- also entail the danger that the produced pseudolabels are highly biased. Surprisingly, given everyday use and the potential for increased bias, weak supervision has not been studied from the point of view of fairness. We begin such a study, starting with the observation that even when a fair model can be built from a dataset with access to ground-truth labels, the corresponding dataset labeled via weak supervision can be arbitrarily unfair. To address this, we propose and empirically validate a model for source unfairness in weak supervision, then introduce a simple counterfactual fairness-based technique that can mitigate these biases. Theoretically, we show that it is possible for our approach to simultaneously improve both accuracy and fairness -- in contrast to standard fairness approaches that suffer from tradeoffs. Empirically, we show that our technique improves accuracy on weak supervision baselines by as much as 32% while reducing demographic parity gap by 82.5%. A simple extension of our method aimed at maximizing performance produces state-of-the-art performance in five out of ten datasets in the WRENCH benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie 等ICML 2021 · 被引用 1,773 次
- No Subclass Left Behind: Fine-Grained Robustness in Coarse-Grained Classification ProblemsNimit Sharad Sohoni, Jared Dunnmon, Geoffrey Angus, Albert Gu 等NeurIPS 2020 · 被引用 316 次
- LIFT: Language-Interfaced Fine-Tuning for Non-language Machine Learning TasksTuan Dinh, Yuchen Zeng, Ruisu Zhang, Ziqian Lin 等NeurIPS 2022 · 被引用 222 次
- Fast and Three-rious: Speeding Up Weak Supervision with Triplet MethodsDaniel Y. Fu, Mayee F. Chen, Frederic Sala, Sarah M. Hooper 等ICML 2020 · 被引用 130 次
相关 Paper
- Towards Harmless Rawlsian Fairness Regardless of Demographic PriorXuanqian Wang, Jing Li, Ivor W. Tsang, Yew Soon OngNeurIPS 2024 · 被引用 3 次
- The Rich Get Richer: Disparate Impact of Semi-Supervised LearningZhaowei Zhu, Tianyi Luo, Yang LiuICLR 2022 · 被引用 44 次
- Debiased Learning from Naturally Imbalanced Pseudo-LabelsXudong Wang, Zhirong Wu, Long Lian, Stella X. YuCVPR 2022 · 被引用 83 次
- Fair Deepfake Detectors Can GeneralizeHarry Cheng, Ming-Hui Liu, Yangyang Guo, Tianyi Wang 等NeurIPS 2025 · 被引用 11 次
- Fair Generative Modeling via Weak SupervisionKristy Choi, Aditya Grover, Trisha Singh, Rui Shu 等ICML 2020 · 被引用 160 次
