A General Framework for Learning from Weak Supervision
Hao Chen, Jindong Wang, Lei Feng, Xiang Li, Yidong Wang, Xing Xie, Masashi Sugiyama, Rita Singh, Bhiksha Raj
Abstract
Weakly supervised learning generally faces challenges in applicability to various scenarios with diverse weak supervision and in scalability due to the complexity of existing algorithms, thereby hindering the practical deployment. This paper introduces a general framework for learning from weak supervision (GLWS) with a novel algorithm. Central to GLWS is an Expectation-Maximization (EM) formulation, adeptly accommodating various weak supervision sources, including instance partial labels, aggregate statistics, pairwise observations, and unlabeled data. We further present an advanced algorithm that significantly simplifies the EM computational demands using a Non-deterministic Finite Automaton (NFA) along with a forward-backward algorithm, which effectively reduces time complexity from quadratic or factorial often required in existing solutions to linear scale. The problem of learning from arbitrary weak supervision is therefore converted to the NFA modeling of them. GLWS not only enhances the scalability of machine learning models but also demonstrates superior performance and versatility across 11 weak supervision scenarios. We hope our work paves the way for further advancements and practical deployment in this field. Code is available at: https://github.com/Hhhhhhao/ General-Framework-Weak-Supervision .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a0f5d92-9bd2-4041-a278-4d1368bf0305Cited by top-tier papers4
- PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning OptimizationYidong Wang, Zhuohao Yu, Wenjin Yao, Zhengran Zeng et al.ICLR 2024 · 368 citations
- Accessible, Realistic, and Fair Evaluation of Positive-Unlabeled Learning AlgorithmsWei Wang, Dong-Dong Wu, Ming Li, Jingxiong Zhang et al.ICLR 2026 · 2 citations
- Weak-to-Strong Generalization with Failure TrajectoriesRuimeng Ye, Zihan Wang, Yang Xiao, Zinan Ling et al.ICLR 2026 · 1 citation
- Robust Label Proportions LearningJueyu Chen, Wantao Wen, Yeqiang Wang, Erliang Lin et al.NeurIPS 2025
Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo LabelingBowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu et al.NeurIPS 2021 · 1,389 citations
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski et al.ICML 2023 · 848 citations
Related papers
- Learning Hyper Label Model for Programmatic Weak SupervisionRenzhi Wu, Shen-En Chen, Jieyu Zhang, Xu ChuICLR 2023 · 2 citations
- Creating Training Sets via Weak Indirect SupervisionJieyu Zhang, Bohan Wang, Xiangchen Song, Yujing Wang et al.ICLR 2022 · 17 citations
- Label Propagation with Weak SupervisionRattana Pukdee, Dylan Sam, Pradeep Kumar Ravikumar, Nina BalcanICLR 2023
- Imprecise Label Learning: A Unified Framework for Learning with Various Imprecise Label ConfigurationsHao Chen, Ankit Shah, Jindong Wang, Ran Tao et al.NeurIPS 2024 · 22 citations
- Disambiguation of Weak Supervision leading to Exponential Convergence ratesVivien A. Cabannes, Francis R. Bach, Alessandro RudiICML 2021 · 6 citations
