Learning from Label Proportions: A Mutual Contamination Framework
Clayton Scott, Jianxin Zhang
Abstract
Learning from label proportions (LLP) is a weakly supervised setting for classification in which unlabeled training instances are grouped into bags, and each bag is annotated with the proportion of each class occurring in that bag. Prior work on LLP has yet to establish a consistent learning procedure, nor does there exist a theoretically justified, general purpose training criterion. In this work we address these two issues by posing LLP in terms of mutual contamination models (MCMs), which have recently been applied successfully to study various other weak supervision settings. In the process, we establish several novel technical results for MCMs, including unbiased losses and generalization error bounds under non-iid sampling plans. We also point out the limitations of a common experimental setting for LLP, and propose a new one based on our MCM framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ae1d8c66-0eab-412d-896e-7f623e48c225Cited by top-tier papers22
- Learning from Label Proportions by Learning with Label NoiseJianxin Zhang, Yutong Wang, Clayton ScottNeurIPS 2022 · 41 citations
- Easy Learning from Label ProportionsRóbert Busa-Fekete, Heejin Choi, Travis Dick, Claudio Gentile et al.NeurIPS 2023 · 24 citations
- Imprecise Label Learning: A Unified Framework for Learning with Various Imprecise Label ConfigurationsHao Chen, Ankit Shah, Jindong Wang, Ran Tao et al.NeurIPS 2024 · 22 citations
- Learnability of Linear Thresholds from Label ProportionsRishi SaketNeurIPS 2021 · 19 citations
- Binary Classification from Multiple Unlabeled Datasets via Surrogate Set ClassificationNan Lu, Shida Lei, Gang Niu, Issei Sato et al.ICML 2021 · 17 citations
Related papers
- MixBag: Bag-Level Data Augmentation for Learning from Label ProportionsTakanori Asanomi, Shinnosuke Matsuo, Daiki Suehiro, Ryoma BiseICCV 2023 · 13 citations
- Forming Auxiliary High-confident Instance-level Loss to Promote Learning from Label ProportionsTianhao Ma, Han Chen, Juncheng Hu, Yungang Zhu et al.CVPR 2025
- Robust Label Proportions LearningJueyu Chen, Wantao Wen, Yeqiang Wang, Erliang Lin et al.NeurIPS 2025
- Dependence and Model Selection in LLP: The Problem of VariantsGabriel Franco, Mark Crovella, Giovanni ComarelaKDD 2023 · 2 citations
- PAC Learning Linear Thresholds from Label ProportionsAnand Brahmbhatt, Rishi Saket, Aravindan RaghuveerNeurIPS 2023 · 12 citations
