Lune

NeurIPS2025Top-tier venue

On the Mechanisms of Weak-to-Strong Generalization: A Theoretical Perspective

Behrad Moniri, Hamed Hassani

2025Year
8Citations
7Top-tier citations

Abstract

Weak-to-strong generalization-where a student model trained on imperfect labels generated by a weaker teacher nonetheless surpasses that teacher-has been widely observed, but the mechanisms that enable it have remained poorly understood. In this paper, through a theoretical analysis of simple models, we uncover three core mechanisms that can drive this phenomenon. First, by analyzing ridge linear regression, we study the interplay between the teacher and student regularization parameters and prove that a student can compensate for a teacher's under-regularization and achieve lower test error. We also analyze the role of the parameterization regime of the models and show that qualitatively different phenomena can happen in different regimes. Second, by analyzing weighted ridge linear regression, we show that a student model with a regularization structure better aligned to the target function, can outperform its teacher. Third, in a nonlinear multi-index learning setting, we demonstrate that a student can learn easy, task-specific features from the teacher while leveraging its own broader pre-training to learn hard-to-learn features that the teacher cannot capture.

Weak-to-strong generalization refers to the phenomenon where a strong (student) model trained on data produced by a weak (teacher) model can sometimes significantly surpass the teacher's performance. This concept was first introduced by Burns et al. [2024], where the authors fine-tuned the GPT-2 model [Radford et al., 2019] (the teacher) for a specific task using ground-truth labels, subsequently employing the fine-tuned model to generate synthetic samples for the same task. These synthetic samples were then used to fine-tune GPT-4 [Achiam et al., 2023] (the student). Remarkably, the fine-tuned student model outperformed its teacher in certain settings despite having access only to the imperfect synthetic data generated by the teacher.

Weak-to-strong generalization is an especially important phenomenon from a practical perspective because of its implications for the emerging question of superalignment [OpenAI, 2023]; i.e., can humans steer models with potentially superhuman capabilities to become aligned to human norms and values [Burns et al., 2024]? Considering the weak model as a proxy for humans, the possibility of the weak-to-strong generalization phenomenon suggests that the answer can be affirmative.

Despite its practical importance, the mechanisms that enable weak-to-strong generalization are still not fully understood. Regularization has empirically been shown to play a critical role in enabling weak-to-strong generalization. However, despite recent theoretical progress demonstrating that regularizing the student is necessary in some settings [Medvedev et al., 2025], the full picture of the effects and the interplay of the regularization of both the student and teacher models in weak-to-strong generalization is still unclear. For example, prior work on weak-to-strong generalization mainly focus on ridgeless regression (see e.g.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 88756eda-6c49-4bab-b08f-d7655ec75e32

Cited by top-tier papers7

Ask how each one uses it

Builds on28

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines