Bayesian WeakS-to-Strong from Text Classification to Generation
Ziyun Cui, Ziyang Zhang, Guangzhi Sun, Wen Wu, Chao Zhang
摘要
Advances in large language models raise the question of how alignment techniques will adapt as models become increasingly complex and humans will only be able to supervise them weakly. Weak-to-Strong mimics such a scenario where weak model supervision attempts to harness the full capabilities of a much stronger model. This work extends Weak-to-Strong to WeakS-to-Strong by exploring an ensemble of weak models which simulate the variability in human opinions. Confidence scores are estimated using a Bayesian approach to guide the WeakSto-Strong generalization. Furthermore, we extend the application of WeakS-to-Strong from text classification tasks to text generation tasks where more advanced strategies are investigated for supervision. Moreover, direct preference optimization is applied to advance the student model's preference learning, beyond the basic learning framework of teacher forcing. Results demonstrate the effectiveness of the proposed approach for the reliability of a strong student model, showing potential for superalignment. 1 2 * Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Weak-to-Strong Generalization under Distribution ShiftsMyeongho Jeon, Jan Sobotka, Suhwan Choi, Maria BrbicNeurIPS 2025 · 被引用 6 次
- Weak-to-Strong Generalization via Bregman Bias–Variance DecompositionGengze Xu, Wei Yao, Ziqiao Wang, Yong LiuICML 2026 · 被引用 4 次
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker 等ICML 2024 · 被引用 443 次
相关 Paper
- Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong GeneralizationWenkai Yang, Shiqi Shen, Guangyao Shen, Wei Yao 等ICLR 2025
- Weak to Strong Generalization for Large Language Models with Multi-capabilitiesYucheng Zhou, Jianbing Shen, Yu ChengICLR 2025
- A transfer learning framework for weak to strong generalizationSeamus Somerstep, Felipe Maia Polo, Moulinath Banerjee, Yaacov Ritov 等ICLR 2025
- How to Mitigate Overfitting in Weak-to-strong Generalization?Junhao Shi, Qinyuan Cheng, Zhaoye Fei, Yining Zheng 等ACL 2025 · 被引用 1 次
- Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned ModelWenhong Zhu, Zhiwei He, Xiaofeng Wang, Pengfei Liu 等ICLR 2025
