High-dimensional Analysis of Knowledge Distillation: Weak-to-Strong Generalization and Scaling Laws
Muhammed Emrullah Ildiz, Halil Alperen Gozeten, Ege Onur Taga, Marco Mondelli, Samet Oymak
摘要
A growing number of machine learning scenarios rely on knowledge distillation where one uses the output of a surrogate model as labels to supervise the training of a target model. In this work, we provide a sharp characterization of this process for ridgeless, high-dimensional regression, under two settings: (i) model shift, where the surrogate model is arbitrary, and (ii) distribution shift, where the surrogate model is the solution of empirical risk minimization with out-of-distribution data. In both cases, we characterize the precise risk of the target model through non-asymptotic bounds in terms of sample size and data distribution under mild conditions. As a consequence, we identify the form of the optimal surrogate model, which reveals the benefits and limitations of discarding weak features in a data-dependent fashion. In the context of weak-to-strong (W2S) generalization, this has the interpretation that (i) W2S training, with the surrogate as the weak model, can provably outperform training with strong labels under the same data budget, but (ii) it is unable to improve the data scaling law. We validate our results on numerical experiments both on ridgeless regression and on neural network architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Preventing Model Collapse Under Overparametrization: Optimal Mixing Ratios for Interpolation Learning and Ridge RegressionAnvit Garg, Sohom Bhattacharya, Pragya SurICLR 2026 · 被引用 9 次
- On the Mechanisms of Weak-to-Strong Generalization: A Theoretical PerspectiveBehrad Moniri, Hamed HassaniNeurIPS 2025 · 被引用 8 次
- Weak-to-Strong Generalization under Distribution ShiftsMyeongho Jeon, Jan Sobotka, Suhwan Choi, Maria BrbicNeurIPS 2025 · 被引用 6 次
- High-dimensional Analysis of Synthetic Data SelectionParham Rezaei, Filip Kovacevic, Francesco Locatello, Marco MondelliICLR 2026 · 被引用 6 次
- Prediction-Powered Semi-Supervised Learning with Online Power TuningNoa Shoham, Ron Dorfman, Shalev Shaer, Kfir Y. Levy 等NeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper22
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker 等ICML 2024 · 被引用 443 次
- Self-Distillation Amplifies Regularization in Hilbert SpaceHossein Mobahi, Mehrdad Farajtabar, Peter L. BartlettNeurIPS 2020 · 被引用 298 次
- Theoretical Analysis of Self-Training with Deep Networks on Unlabeled DataColin Wei, Kendrick Shen, Yining Chen, Tengyu MaICLR 2021 · 被引用 261 次
相关 Paper
- Improved Scaling Laws via Weak-to-Strong Generalization in Random Features Ridge RegressionDiyuan Wu, Lehan Chen, Theodor Misiakiewicz, Marco MondelliICML 2026
- Discrepancies are Virtue: Weak-to-Strong Generalization through Lens of Intrinsic DimensionYijun Dong, Yicheng Li, Yunai Li, Jason D. Lee 等ICML 2025
- Optimal Unconstrained Self-Distillation in Ridge Regression: Strict Improvements, Precise Asymptotics, and One-Shot TuningHien Dang, Pratik Patil, Alessandro RinaldoICML 2026 · 被引用 1 次
- Quantifying Cross-Domain Knowledge Distillation in the Presence of Domain ShiftXiangchao Li, Xiao Han, Qing Yang, Xin TongICML 2026
- Semi-Supervised Learning with Noisy Proxy Covariates: Generalization Bounds and Distribution RegressionKwangho Kim, Jisu KimICML 2026
