Lune

NeurIPS2025顶会

On the Mechanisms of Weak-to-Strong Generalization: A Theoretical Perspective

Behrad Moniri, Hamed Hassani

2025年份
8被引次数
7顶会引用

摘要

Weak-to-strong generalization-where a student model trained on imperfect labels generated by a weaker teacher nonetheless surpasses that teacher-has been widely observed, but the mechanisms that enable it have remained poorly understood. In this paper, through a theoretical analysis of simple models, we uncover three core mechanisms that can drive this phenomenon. First, by analyzing ridge linear regression, we study the interplay between the teacher and student regularization parameters and prove that a student can compensate for a teacher's under-regularization and achieve lower test error. We also analyze the role of the parameterization regime of the models and show that qualitatively different phenomena can happen in different regimes. Second, by analyzing weighted ridge linear regression, we show that a student model with a regularization structure better aligned to the target function, can outperform its teacher. Third, in a nonlinear multi-index learning setting, we demonstrate that a student can learn easy, task-specific features from the teacher while leveraging its own broader pre-training to learn hard-to-learn features that the teacher cannot capture.

Weak-to-strong generalization refers to the phenomenon where a strong (student) model trained on data produced by a weak (teacher) model can sometimes significantly surpass the teacher's performance. This concept was first introduced by Burns et al. [2024], where the authors fine-tuned the GPT-2 model [Radford et al., 2019] (the teacher) for a specific task using ground-truth labels, subsequently employing the fine-tuned model to generate synthetic samples for the same task. These synthetic samples were then used to fine-tune GPT-4 [Achiam et al., 2023] (the student). Remarkably, the fine-tuned student model outperformed its teacher in certain settings despite having access only to the imperfect synthetic data generated by the teacher.

Weak-to-strong generalization is an especially important phenomenon from a practical perspective because of its implications for the emerging question of superalignment [OpenAI, 2023]; i.e., can humans steer models with potentially superhuman capabilities to become aligned to human norms and values [Burns et al., 2024]? Considering the weak model as a proxy for humans, the possibility of the weak-to-strong generalization phenomenon suggests that the answer can be affirmative.

Despite its practical importance, the mechanisms that enable weak-to-strong generalization are still not fully understood. Regularization has empirically been shown to play a critical role in enabling weak-to-strong generalization. However, despite recent theoretical progress demonstrating that regularizing the student is necessary in some settings [Medvedev et al., 2025], the full picture of the effects and the interplay of the regularization of both the student and teacher models in weak-to-strong generalization is still unclear. For example, prior work on weak-to-strong generalization mainly focus on ridgeless regression (see e.g.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper7

问问它们各自怎么用它

它引用的顶会 Paper28

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖