A transfer learning framework for weak to strong generalization
Seamus Somerstep, Felipe Maia Polo, Moulinath Banerjee, Yaacov Ritov, Mikhail Yurochkin, Yuekai Sun
摘要
Modern large language model (LLM) alignment techniques rely on human feedback, but it is unclear whether these techniques fundamentally limit the capabilities of aligned LLMs. In particular, it is unknown if it is possible to align (stronger) LLMs with superhuman capabilities with (weaker) human feedback without degrading their capabilities. This is an instance of the weak-to-strong generalization problem: using feedback from a weaker (less capable) model to train a stronger (more capable) model. We prove that weak-to-strong generalization is possible by eliciting latent knowledge from pre-trained LLMs. In particular, we cast the weakto-strong generalization problem as a transfer learning problem in which we wish to transfer a latent concept prior from a weak model to a strong pre-trained model. We prove that a naive fine-tuning approach suffers from fundamental limitations, but an alternative refinement-based approach suggested by the problem structure provably overcomes the limitations of fine-tuning. Finally, we demonstrate the practical applicability of the refinement approach in multiple LLM alignment tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Weak-to-Strong Generalization via Bregman Bias–Variance DecompositionGengze Xu, Wei Yao, Ziqiao Wang, Yong LiuICML 2026 · 被引用 4 次
- Contrastive Weak-to-Strong GeneralizationHoucheng Jiang, Junfeng Fang, Jiaxin Wu, Tianyu Zhang 等ICML 2026 · 被引用 2 次
- Improved Scaling Laws via Weak-to-Strong Generalization in Random Features Ridge RegressionDiyuan Wu, Lehan Chen, Theodor Misiakiewicz, Marco MondelliICML 2026
- Trust Functions: Near Lossless Weak-to-Strong Generalization by Learning to Trust the Weak TeacherArda Uzunoglu, Alvin Zhang, Daniel KhashabiICML 2026
- GRAD: Generalizing RAG Adaptation with DecodingYoungwon Lee, Seung-won Hwang, Zhewei Yao, Yuxiong HeACL 2026
它引用的顶会 Paper26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
- DivideMix: Learning with Noisy Labels as Semi-supervised LearningJunnan Li, Richard Socher, Steven C. H. HoiICLR 2020 · 被引用 1,326 次
相关 Paper
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker 等ICML 2024 · 被引用 443 次
- How to Mitigate Overfitting in Weak-to-strong Generalization?Junhao Shi, Qinyuan Cheng, Zhaoye Fei, Yining Zheng 等ACL 2025 · 被引用 1 次
- Weak to Strong Generalization for Large Language Models with Multi-capabilitiesYucheng Zhou, Jianbing Shen, Yu ChengICLR 2025
- Self-Refine Instruction-Tuning for Aligning Reasoning in Language ModelsLeonardo Ranaldi, André FreitasEMNLP 2024 · 被引用 3 次
- Synergistic Weak-Strong Collaboration by Aligning PreferencesYizhu Jiao, Xuchao Zhang, Zhaoyang Wang, Yubo Ma 等ACL 2025
