Representations Shape Weak-to-Strong Generalization: Theoretical Insights and Empirical Predictions
Yihao Xue, Jiping Li, Baharan Mirzasoleiman
Abstract
Weak-to-Strong Generalization (W2SG), where a weak model supervises a stronger one, serves as an important analogy for understanding how humans might guide superhuman intelligence in the future. Promising empirical results revealed that a strong model can surpass its weak supervisor. While recent work has offered theoretical insights into this phenomenon, a clear understanding of the interactions between weak and strong models that drive W2SG remains elusive. We investigate W2SG through a theoretical lens and show that it can be characterized using kernels derived from the principal components of weak and strong models' internal representations. These kernels can be used to define a space that, at a high level, captures what the weak model is unable to learn but is learnable by the strong model. The projection of labels onto this space quantifies how much the strong model falls short of its full potential due to weak supervision. This characterization also provides insights into how certain errors in weak supervision can be corrected by the strong model, regardless of overfitting. Our theory has significant practical implications, providing a representationbased metric that predicts W2SG performance trends without requiring labels, as shown in experiments on molecular predictions with transformers and 5 NLP tasks involving 52 LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5cf6f73f-e962-428b-bc14-d520ebbb2344Cited by top-tier papers6
- On the Mechanisms of Weak-to-Strong Generalization: A Theoretical PerspectiveBehrad Moniri, Hamed HassaniNeurIPS 2025 · 8 citations
- Does Weak-to-strong Generalization Happen under Spurious Correlations?Chenruo Liu, Yijun Dong, Qi LeiICLR 2026 · 1 citation
- DC-W2S: Dual-Consensus Weak-to-Strong Training for Reliable Process Reward Modeling in Biological ReasoningChi-Min Chan, Ehsan Hajiramezanali, Xiner Li, Edward De Brouwer et al.ICML 2026 · 1 citation
- Why Self-Distillation Helps and Hurts: Denoising vs. Signal ForgettingMingqi Wu, Archer Yang, Qiang SunICML 2026
- Improved Scaling Laws via Weak-to-Strong Generalization in Random Features Ridge RegressionDiyuan Wu, Lehan Chen, Theodor Misiakiewicz, Marco MondelliICML 2026
Builds on16
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
- Toward Understanding the Feature Learning Process of Self-supervised Contrastive LearningZixin Wen, Yuanzhi LiICML 2021 · 162 citations
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 151 citations
- A Kernel-Based View of Language Model Fine-TuningSadhika Malladi, Alexander Wettig, Dingli Yu, Danqi Chen et al.ICML 2023 · 111 citations
- Implicit Regularization of Random Feature ModelsArthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler et al.ICML 2020 · 83 citations
Related papers
- Quantifying the Gain in Weak-to-Strong GeneralizationMoses Charikar, Chirag Pabbaraju, Kirankumar ShiragurNeurIPS 2024 · 42 citations
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker et al.ICML 2024 · 443 citations
- How to Mitigate Overfitting in Weak-to-strong Generalization?Junhao Shi, Qinyuan Cheng, Zhaoye Fei, Yining Zheng et al.ACL 2025 · 1 citation
- Discrepancies are Virtue: Weak-to-Strong Generalization through Lens of Intrinsic DimensionYijun Dong, Yicheng Li, Yunai Li, Jason D. Lee et al.ICML 2025
- Weak-to-Strong Generalization via Bregman Bias–Variance DecompositionGengze Xu, Wei Yao, Ziqiao Wang, Yong LiuICML 2026 · 4 citations
