Quantifying the Gain in Weak-to-Strong Generalization
Moses Charikar, Chirag Pabbaraju, Kirankumar Shiragur
摘要
Recent advances in large language models have shown capabilities that are extraordinary and near-superhuman. These models operate with such complexity that reliably evaluating and aligning them proves challenging for humans. This leads to the natural question: can guidance from weak models (like humans) adequately direct the capabilities of strong models? In a recent and somewhat surprising work, Burns et al. (2023) empirically demonstrated that when strong models (like GPT-4) are finetuned using labels generated by weak supervisors (like GPT-2), the strong models outperform their weaker counterparts -- a phenomenon they term weak-to-strong generalization. In this work, we present a theoretical framework for understanding weak-to-strong generalization. Specifically, we show that the improvement in performance achieved by strong models over their weaker counterparts is quantified by the misfit error incurred by the strong model on labels generated by the weaker model. Our theory reveals several curious algorithmic insights. For instance, we can predict the amount by which the strong model will improve over the weak model, and also choose among different weak models to train the strong model, based on its misfit error. We validate our theoretical findings through various empirical assessments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Selective Preference Optimization via Token-Level Reward Function EstimationKailai Yang, Zhiwei Liu, Qianqian Xie, Jimin Huang 等EMNLP 2025 · 被引用 18 次
- On the Mechanisms of Weak-to-Strong Generalization: A Theoretical PerspectiveBehrad Moniri, Hamed HassaniNeurIPS 2025 · 被引用 8 次
- Weak-to-Strong Generalization under Distribution ShiftsMyeongho Jeon, Jan Sobotka, Suhwan Choi, Maria BrbicNeurIPS 2025 · 被引用 6 次
- From Linear to Nonlinear: Provable Weak-to-Strong Generalization through Feature LearningJunsoo Oh, Jerry Song, Chulhee YunNeurIPS 2025 · 被引用 5 次
- Weak-to-Strong Generalization via Bregman Bias–Variance DecompositionGengze Xu, Wei Yao, Ziqiao Wang, Yong LiuICML 2026 · 被引用 4 次
它引用的顶会 Paper6
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma 等ICLR 2022 · 被引用 911 次
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker 等ICML 2024 · 被引用 443 次
- Self-Distillation Amplifies Regularization in Hilbert SpaceHossein Mobahi, Mehrdad Farajtabar, Peter L. BartlettNeurIPS 2020 · 被引用 298 次
- Theoretical Analysis of Self-Training with Deep Networks on Unlabeled DataColin Wei, Kendrick Shen, Yining Chen, Tengyu MaICLR 2021 · 被引用 261 次
- Easy-to-Hard Generalization: Scalable Alignment Beyond Human SupervisionZhiqing Sun, Longhui Yu, Yikang Shen, Weiyang Liu 等NeurIPS 2024 · 被引用 125 次
相关 Paper
- Weak-to-Strong Generalization Even in Random Feature Networks, ProvablyMarko Medvedev, Kaifeng Lyu, Dingli Yu, Sanjeev Arora 等ICML 2025
- How to Mitigate Overfitting in Weak-to-strong Generalization?Junhao Shi, Qinyuan Cheng, Zhaoye Fei, Yining Zheng 等ACL 2025 · 被引用 1 次
- Provable weak-to-strong generalization via benign overfittingDavid Xing Wu, Anant SahaiICLR 2025
- Relating Misfit to Gain in Weak-to-Strong Generalization Beyond the Squared LossAbhijeet Mulgund, Chirag PabbarajuICML 2025
- A transfer learning framework for weak to strong generalizationSeamus Somerstep, Felipe Maia Polo, Moulinath Banerjee, Yaacov Ritov 等ICLR 2025
