Theoretical Analysis of Weak-to-Strong Generalization
Hunter Lang, David A. Sontag, Aravindan Vijayaraghavan
摘要
Strong student models can learn from weaker teachers: when trained on the predictions of a weaker model, a strong pretrained student can learn to correct the weak model's errors and generalize to examples where the teacher is not confident, even when these examples are excluded from training. This enables learning from cheap, incomplete, and possibly incorrect label information, such as coarse logical rules or the generations of a language model. We show that existing weak supervision theory fails to account for both of these effects, which we call pseudolabel correction and coverage expansion, respectively. We give a new bound based on expansion properties of the data distribution and student hypothesis class that directly accounts for pseudolabel correction and coverage expansion. Our bounds capture the intuition that weak-to-strong generalization occurs when the strong model is unable to fit the mistakes of the weak teacher without incurring additional error. We show that these expansion properties can be checked from finite data and give empirical evidence that they hold in practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Scaling Laws For Scalable OversightJoshua Engels, David D. Baek, Subhash Kantamneni, Max TegmarkNeurIPS 2025 · 被引用 22 次
- Selective Preference Optimization via Token-Level Reward Function EstimationKailai Yang, Zhiwei Liu, Qianqian Xie, Jimin Huang 等EMNLP 2025 · 被引用 18 次
- Towards Acyclic Preference Evaluation of Language Models via Multiple EvaluatorsZhengyu Hu, Jieyu Zhang, Zhihan Xiong, Alexander Ratner 等AAAI 2026 · 被引用 14 次
- Robust SuperAlignment: Weak-to-Strong Robustness Generalization for Vision-Language ModelsJunhao Dong, Cong Zhang, Xinghua Qu, Zejun Ma 等NeurIPS 2025 · 被引用 7 次
- Weak-to-Strong Generalization under Distribution ShiftsMyeongho Jeon, Jan Sobotka, Suhwan Choi, Maria BrbicNeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper26
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker 等ICML 2024 · 被引用 443 次
- Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive LossJeff Z. HaoChen, Colin Wei, Adrien Gaidon, Tengyu MaNeurIPS 2021 · 被引用 425 次
- Does Knowledge Distillation Really Work?Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi 等NeurIPS 2021 · 被引用 318 次
- Large language models are few-shot clinical information extractorsMonica Agrawal, Stefan Hegselmann, Hunter Lang, Yoon Kim 等EMNLP 2022 · 被引用 285 次
相关 Paper
- Provable weak-to-strong generalization via benign overfittingDavid Xing Wu, Anant SahaiICLR 2025
- Does Weak-to-strong Generalization Happen under Spurious Correlations?Chenruo Liu, Yijun Dong, Qi LeiICLR 2026 · 被引用 1 次
- Discrepancies are Virtue: Weak-to-Strong Generalization through Lens of Intrinsic DimensionYijun Dong, Yicheng Li, Yunai Li, Jason D. Lee 等ICML 2025
- Disentangling Latent Shifts of In-Context Learning with Weak SupervisionJosip Jukic, Jan SnajderNeurIPS 2025 · 被引用 3 次
- Trust Functions: Near Lossless Weak-to-Strong Generalization by Learning to Trust the Weak TeacherArda Uzunoglu, Alvin Zhang, Daniel KhashabiICML 2026
