Theoretical Analysis of Weak-to-Strong Generalization
Hunter Lang, David A. Sontag, Aravindan Vijayaraghavan
Abstract
Strong student models can learn from weaker teachers: when trained on the predictions of a weaker model, a strong pretrained student can learn to correct the weak model's errors and generalize to examples where the teacher is not confident, even when these examples are excluded from training. This enables learning from cheap, incomplete, and possibly incorrect label information, such as coarse logical rules or the generations of a language model. We show that existing weak supervision theory fails to account for both of these effects, which we call pseudolabel correction and coverage expansion, respectively. We give a new bound based on expansion properties of the data distribution and student hypothesis class that directly accounts for pseudolabel correction and coverage expansion. Our bounds capture the intuition that weak-to-strong generalization occurs when the strong model is unable to fit the mistakes of the weak teacher without incurring additional error. We show that these expansion properties can be checked from finite data and give empirical evidence that they hold in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f89ea0c3-151d-4fe2-a96d-d235c1e153e2Cited by top-tier papers27
- Scaling Laws For Scalable OversightJoshua Engels, David D. Baek, Subhash Kantamneni, Max TegmarkNeurIPS 2025 · 22 citations
- Selective Preference Optimization via Token-Level Reward Function EstimationKailai Yang, Zhiwei Liu, Qianqian Xie, Jimin Huang et al.EMNLP 2025 · 18 citations
- Towards Acyclic Preference Evaluation of Language Models via Multiple EvaluatorsZhengyu Hu, Jieyu Zhang, Zhihan Xiong, Alexander Ratner et al.AAAI 2026 · 14 citations
- Robust SuperAlignment: Weak-to-Strong Robustness Generalization for Vision-Language ModelsJunhao Dong, Cong Zhang, Xinghua Qu, Zejun Ma et al.NeurIPS 2025 · 7 citations
- Weak-to-Strong Generalization under Distribution ShiftsMyeongho Jeon, Jan Sobotka, Suhwan Choi, Maria BrbicNeurIPS 2025 · 6 citations
Builds on26
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker et al.ICML 2024 · 443 citations
- Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive LossJeff Z. HaoChen, Colin Wei, Adrien Gaidon, Tengyu MaNeurIPS 2021 · 425 citations
- Does Knowledge Distillation Really Work?Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi et al.NeurIPS 2021 · 318 citations
- Large language models are few-shot clinical information extractorsMonica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim et al.EMNLP 2022 · 285 citations
Related papers
- Provable weak-to-strong generalization via benign overfittingDavid Xing Wu, Anant SahaiICLR 2025
- Does Weak-to-strong Generalization Happen under Spurious Correlations?Chenruo Liu, Yijun Dong, Qi LeiICLR 2026 · 1 citation
- Discrepancies are Virtue: Weak-to-Strong Generalization through Lens of Intrinsic DimensionYijun Dong, Yicheng Li, Yunai Li, Jason D. Lee et al.ICML 2025
- Disentangling Latent Shifts of In-Context Learning with Weak SupervisionJosip Jukic, Jan SnajderNeurIPS 2025 · 3 citations
- Trust Functions: Near Lossless Weak-to-Strong Generalization by Learning to Trust the Weak TeacherArda Uzunoglu, Alvin Zhang, Daniel KhashabiICML 2026
