Nearly Optimal Sample Complexity for Learning with Label Proportions
Róbert Istvan Busa-Fekete, Travis Dick, Claudio Gentile, Haim Kaplan, Tomer Koren, Uri Stemmer
Abstract
We investigate Learning from Label Proportions (LLP), a partial information setting where examples in a training set are grouped into bags, and only aggregate label values in each bag are available. Despite the partial observability, the goal is still to achieve small regret at the level of individual examples. We give results on the sample complexity of LLP under square loss, showing that our sample complexity is essentially optimal. From an algorithmic viewpoint, we rely on carefully designed variants of Empirical Risk Minimization, and Stochastic Gradient Descent algorithms, combined with ad hoc variance reduction techniques. On one hand, our theoretical results improve in important ways on the existing literature on LLP, specifically in the way the sample complexity depends on the bag size. On the other hand, we validate our algorithmic solutions on several datasets, demonstrating improved empirical performance (better accuracy for less samples) against recent baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Optimal Learning from Label Proportions with General Loss FunctionsLorne Applebaum, Travis Dick, Claudio Gentile, Haim Kaplan et al.ICML 2026 · 1 citation
- Learning from Label Proportions via Proportional Value ClassificationTianhao Ma, Wei Wang, Ximing Li, Gang Niu et al.ICLR 2026
Builds on8
- Learning from Label Proportions by Learning with Label NoiseJianxin Zhang, Yutong Wang, Clayton ScottNeurIPS 2022 · 41 citations
- Easy Learning from Label ProportionsRóbert Busa-Fekete, Heejin Choi, Travis Dick, Claudio Gentile et al.NeurIPS 2023 · 24 citations
- Learnability of Linear Thresholds from Label ProportionsRishi SaketNeurIPS 2021 · 19 citations
- Binary Classification from Multiple Unlabeled Datasets via Surrogate Set ClassificationNan Lu, Shida Lei, Gang Niu, Issei Sato et al.ICML 2021 · 17 citations
- Algorithms and Hardness for Learning Linear Thresholds from Label ProportionsRishi SaketNeurIPS 2022 · 15 citations
Related papers
- Learning from Label Proportions: Bootstrapping Supervised Learners via Belief PropagationShreyas Havaldar, Navodita Sharma, Shubhi Sareen, Karthikeyan Shanmugam et al.ICLR 2024 · 5 citations
- Dependence and Model Selection in LLP: The Problem of VariantsGabriel Franco, Mark Crovella, Giovanni ComarelaKDD 2023 · 2 citations
- PAC Learning Linear Thresholds from Label ProportionsAnand Brahmbhatt, Rishi Saket, Aravindan RaghuveerNeurIPS 2023 · 12 citations
- MixBag: Bag-Level Data Augmentation for Learning from Label ProportionsTakanori Asanomi, Shinnosuke Matsuo, Daiki Suehiro, Ryoma BiseICCV 2023 · 13 citations
- Learning from Label Proportions: A Mutual Contamination FrameworkClayton Scott, Jianxin ZhangNeurIPS 2020 · 12 citations
