On the Mechanisms of Weak-to-Strong Generalization: A Theoretical Perspective
Behrad Moniri, Hamed Hassani
Abstract
Weak-to-strong generalization-where a student model trained on imperfect labels generated by a weaker teacher nonetheless surpasses that teacher-has been widely observed, but the mechanisms that enable it have remained poorly understood. In this paper, through a theoretical analysis of simple models, we uncover three core mechanisms that can drive this phenomenon. First, by analyzing ridge linear regression, we study the interplay between the teacher and student regularization parameters and prove that a student can compensate for a teacher's under-regularization and achieve lower test error. We also analyze the role of the parameterization regime of the models and show that qualitatively different phenomena can happen in different regimes. Second, by analyzing weighted ridge linear regression, we show that a student model with a regularization structure better aligned to the target function, can outperform its teacher. Third, in a nonlinear multi-index learning setting, we demonstrate that a student can learn easy, task-specific features from the teacher while leveraging its own broader pre-training to learn hard-to-learn features that the teacher cannot capture.
Weak-to-strong generalization refers to the phenomenon where a strong (student) model trained on data produced by a weak (teacher) model can sometimes significantly surpass the teacher's performance. This concept was first introduced by Burns et al. [2024], where the authors fine-tuned the GPT-2 model [Radford et al., 2019] (the teacher) for a specific task using ground-truth labels, subsequently employing the fine-tuned model to generate synthetic samples for the same task. These synthetic samples were then used to fine-tune GPT-4 [Achiam et al., 2023] (the student). Remarkably, the fine-tuned student model outperformed its teacher in certain settings despite having access only to the imperfect synthetic data generated by the teacher.
Weak-to-strong generalization is an especially important phenomenon from a practical perspective because of its implications for the emerging question of superalignment [OpenAI, 2023]; i.e., can humans steer models with potentially superhuman capabilities to become aligned to human norms and values [Burns et al., 2024]? Considering the weak model as a proxy for humans, the possibility of the weak-to-strong generalization phenomenon suggests that the answer can be affirmative.
Despite its practical importance, the mechanisms that enable weak-to-strong generalization are still not fully understood. Regularization has empirically been shown to play a critical role in enabling weak-to-strong generalization. However, despite recent theoretical progress demonstrating that regularizing the student is necessary in some settings [Medvedev et al., 2025], the full picture of the effects and the interplay of the regularization of both the student and teacher models in weak-to-strong generalization is still unclear. For example, prior work on weak-to-strong generalization mainly focus on ridgeless regression (see e.g.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 88756eda-6c49-4bab-b08f-d7655ec75e32Cited by top-tier papers7
- Weak-to-Strong Generalization via Bregman Bias–Variance DecompositionGengze Xu, Wei Yao, Ziqiao Wang, Yong LiuICML 2026 · 4 citations
- Optimal Unconstrained Self-Distillation in Ridge Regression: Strict Improvements, Precise Asymptotics, and One-Shot TuningHien Dang, Pratik Patil, Alessandro RinaldoICML 2026 · 1 citation
- Does Weak-to-strong Generalization Happen under Spurious Correlations?Chenruo Liu, Yijun Dong, Qi LeiICLR 2026 · 1 citation
- Why Self-Distillation Helps and Hurts: Denoising vs. Signal ForgettingMingqi Wu, Archer Yang, Qiang SunICML 2026
- Improved Scaling Laws via Weak-to-Strong Generalization in Random Features Ridge RegressionDiyuan Wu, Lehan Chen, Theodor Misiakiewicz, Marco MondelliICML 2026
Builds on28
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker et al.ICML 2024 · 443 citations
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 242 citations
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang et al.NeurIPS 2022 · 173 citations
- Optimal Regularization can Mitigate Double DescentPreetum Nakkiran, Prayaag Venkat, Sham M. Kakade, Tengyu MaICLR 2021 · 148 citations
- A Kernel-Based View of Language Model Fine-TuningSadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen et al.ICML 2023 · 111 citations
Related papers
- Weak-to-Strong Generalization Even in Random Feature Networks, ProvablyMarko Medvedev, Kaifeng Lyu, Dingli Yu, Sanjeev Arora et al.ICML 2025
- Quantifying the Gain in Weak-to-Strong GeneralizationMoses Charikar, Chirag Pabbaraju, Kirankumar ShiragurNeurIPS 2024 · 42 citations
- Provable weak-to-strong generalization via benign overfittingDavid Xing Wu, Anant SahaiICLR 2025
- From Linear to Nonlinear: Provable Weak-to-Strong Generalization through Feature LearningJunsoo Oh, Jerry Song, Chulhee YunNeurIPS 2025 · 5 citations
- Discrepancies are Virtue: Weak-to-Strong Generalization through Lens of Intrinsic DimensionYijun Dong, Yicheng Li, Yunai Li, Jason D. Lee et al.ICML 2025
