Mitigating Source Bias for Fairer Weak Supervision
Changho Shin, Sonia Cromp, Dyah Adila, Frederic Sala
Abstract
Weak supervision enables efficient development of training sets by reducing the need for ground truth labels. However, the techniques that make weak supervision attractive -- such as integrating any source of signal to estimate unknown labels -- also entail the danger that the produced pseudolabels are highly biased. Surprisingly, given everyday use and the potential for increased bias, weak supervision has not been studied from the point of view of fairness. We begin such a study, starting with the observation that even when a fair model can be built from a dataset with access to ground-truth labels, the corresponding dataset labeled via weak supervision can be arbitrarily unfair. To address this, we propose and empirically validate a model for source unfairness in weak supervision, then introduce a simple counterfactual fairness-based technique that can mitigate these biases. Theoretically, we show that it is possible for our approach to simultaneously improve both accuracy and fairness -- in contrast to standard fairness approaches that suffer from tradeoffs. Empirically, we show that our technique improves accuracy on weak supervision baselines by as much as 32% while reducing demographic parity gap by 82.5%. A simple extension of our method aimed at maximizing performance produces state-of-the-art performance in five out of ten datasets in the WRENCH benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 431c4903-5f6a-4ff1-9a36-03c026c41faaCited by top-tier papers1
Ask how each one uses itBuilds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- No Subclass Left Behind: Fine-Grained Robustness in Coarse-Grained Classification ProblemsNimit Sharad Sohoni, Jared Dunnmon, Geoffrey Angus, Albert Gu et al.NeurIPS 2020 · 316 citations
- LIFT: Language-Interfaced Fine-Tuning for Non-language Machine Learning TasksTuan Dinh, Yuchen Zeng, Ruisu Zhang, Ziqian Lin et al.NeurIPS 2022 · 222 citations
- Fast and Three-rious: Speeding Up Weak Supervision with Triplet MethodsDaniel Y. Fu, Mayee F. Chen, Frederic Sala, Sarah M. Hooper et al.ICML 2020 · 130 citations
Related papers
- Towards Harmless Rawlsian Fairness Regardless of Demographic PriorXuanqian Wang, Jing Li, Ivor W. Tsang, Yew Soon OngNeurIPS 2024 · 3 citations
- The Rich Get Richer: Disparate Impact of Semi-Supervised LearningZhaowei Zhu, Tianyi Luo, Yang LiuICLR 2022 · 44 citations
- Debiased Learning from Naturally Imbalanced Pseudo-LabelsXudong Wang, Zhirong Wu, Long Lian, Stella X. YuCVPR 2022 · 83 citations
- Fair Deepfake Detectors Can GeneralizeHarry Cheng, Ming-Hui Liu, Yangyang Guo, Tianyi Wang et al.NeurIPS 2025 · 11 citations
- Fair Generative Modeling via Weak SupervisionKristy Choi, Aditya Grover, Trisha Singh, Rui Shu et al.ICML 2020 · 160 citations
