Easy Learning from Label Proportions
Róbert Busa-Fekete, Heejin Choi, Travis Dick, Claudio Gentile, Andrés Muñoz Medina
Abstract
We consider the problem of Learning from Label Proportions (LLP), a weakly supervised classification setup where instances are grouped into "bags", and only the frequency of class labels at each bag is available. Albeit, the objective of the learner is to achieve low task loss at an individual instance level. Here we propose EASYLLP: a flexible and simple-to-implement debiasing approach based on aggregate labels, which operates on arbitrary loss functions. Our technique allows us to accurately estimate the expected loss of an arbitrary model at an individual level. We showcase the flexibility of our approach by applying it to popular learning frameworks, like Empirical Risk Minimization (ERM) and Stochastic Gradient Descent (SGD) with provable guarantees on instance level performance. More concretely, we exhibit a variance reduction technique that makes the quality of LLP learning deteriorate only by a factor of k (k being bag size) in both ERM and SGD setups, as compared to full supervision. Finally, we validate our theoretical results on multiple datasets demonstrating our algorithm performs as well or better than previous LLP approaches in spite of its simplicity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2046359-4f8a-4396-abb7-6b68174ef497Cited by top-tier papers11
- PriorBoost: An Adaptive Algorithm for Learning from Aggregate ResponsesAdel Javanmard, Matthew Fahrbach, Vahab MirrokniICML 2024 · 6 citations
- Learning from Label Proportions: Bootstrapping Supervised Learners via Belief PropagationShreyas Havaldar, Navodita Sharma, Shubhi Sareen, Karthikeyan Shanmugam et al.ICLR 2024 · 5 citations
- Learning from Aggregate responses: Instance Level versus Bag Level Loss FunctionsAdel Javanmard, Lin Chen, Vahab Mirrokni, Ashwinkumar Badanidiyuru et al.ICLR 2024 · 3 citations
- Auditing Privacy Mechanisms via Label Inference AttacksRóbert Busa-Fekete, Travis Dick, Claudio Gentile, Andrés Muñoz Medina et al.NeurIPS 2024 · 3 citations
- A Closer Look to Positive-Unlabeled Learning from Fine-grained Perspectives: An Empirical StudyYuanchao Dai, Zhengzhang Hou, Changchun Li, Yuanbo Xu et al.NeurIPS 2025 · 2 citations
Builds on5
- Learning from Label Proportions by Learning with Label NoiseJianxin Zhang, Yutong Wang, Clayton ScottNeurIPS 2022 · 41 citations
- Learnability of Linear Thresholds from Label ProportionsRishi SaketNeurIPS 2021 · 19 citations
- Binary Classification from Multiple Unlabeled Datasets via Surrogate Set ClassificationNan Lu, Shida Lei, Gang Niu, Issei Sato et al.ICML 2021 · 17 citations
- Algorithms and Hardness for Learning Linear Thresholds from Label ProportionsRishi SaketNeurIPS 2022 · 15 citations
- Learning from Label Proportions: A Mutual Contamination FrameworkClayton Scott, Jianxin ZhangNeurIPS 2020 · 12 citations
Related papers
- Optimal Learning from Label Proportions with General Loss FunctionsLorne Applebaum, Travis Dick, Claudio Gentile, Haim Kaplan et al.ICML 2026 · 1 citation
- Nearly Optimal Sample Complexity for Learning with Label ProportionsRóbert Istvan Busa-Fekete, Travis Dick, Claudio Gentile, Haim Kaplan et al.ICML 2025
- MixBag: Bag-Level Data Augmentation for Learning from Label ProportionsTakanori Asanomi, Shinnosuke Matsuo, Daiki Suehiro, Ryoma BiseICCV 2023 · 13 citations
- Forming Auxiliary High-confident Instance-level Loss to Promote Learning from Label ProportionsTianhao Ma, Han Chen, Juncheng Hu, Yungang Zhu et al.CVPR 2025
- Learning from Label Proportions via Proportional Value ClassificationTianhao Ma, Wei Wang, Ximing Li, Gang Niu et al.ICLR 2026
