Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization
Wenkai Yang, Shiqi Shen, Guangyao Shen, Wei Yao, Yong Liu, Gong Zhi, Yankai Lin, Ji-Rong Wen
摘要
Superalignment, where humans act as weak supervisors for superhuman models, has become a crucial problem with the rapid development of Large Language Models (LLMs). Recent work has preliminarily studied this problem by using weak models to supervise strong models, and discovered that weakly supervised strong students can consistently outperform weak teachers towards the alignment target, leading to a weak-to-strong generalization phenomenon. However, we are concerned that behind such a promising phenomenon, whether there exists an issue of weak-to-strong deception, where strong models deceive weak models by exhibiting well-aligned in areas known to weak models but producing misaligned behaviors in cases weak models do not know. We take an initial step towards exploring this security issue in a specific but realistic multi-objective alignment case, where there may be some alignment targets conflicting with each other (e.g., helpfulness v.s. harmlessness). We aim to explore whether, in such cases, strong models might deliberately make mistakes in areas known to them but unknown to weak models within one alignment dimension, in exchange for a higher reward in another dimension. Through extensive experiments in both the reward modeling and preference optimization scenarios, we find: (1) The weak-to-strong deception phenomenon exists across all settings. (2) The deception intensifies as the capability gap between weak and strong models increases. (3) Bootstrapping with an intermediate model can mitigate the deception to some extent, though its effectiveness remains limited. Our work highlights the urgent need to pay more attention to the true reliability of superalignment. 1 * Corresponding Author 1 Code is available at https://github.com/RUCBM/weak-to-strong-deception . Human Superhuman Model (a) Superalignment supervise Weak Model Strong Model supervise (b) Analogous Problem Weak Model Strong Model Strong Model W ea k K n o w n W ea k U n k n o w n St ro ng K no w n St ro ng U nk no w n W ea k K n o w n W ea k U n k n o w n St ro ng K no w n St ro ng U nk no w n s u p e r v is e s u p e r v is e (c) Weak-to-Strong Generalization (d) Weak-to-Strong Deception generate weak supervision data W ea k K n o w n W ea k U n k n o w n St ro ng K no w n St ro ng U nk no w n another conflicting alignment target Generalizing well to areas unknow to the weak model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Weak-to-Strong Search: Align Large Language Models via Searching over Small Language ModelsZhanhui Zhou, Zhixuan Liu, Jie Liu, Zhichen Dong 等NeurIPS 2024 · 被引用 50 次
- Contrastive Weak-to-Strong GeneralizationHoucheng Jiang, Junfeng Fang, Jiaxin Wu, Tianyu Zhang 等ICML 2026 · 被引用 2 次
- Does Weak-to-strong Generalization Happen under Spurious Correlations?Chenruo Liu, Yijun Dong, Qi LeiICLR 2026 · 被引用 1 次
- Weak-to-Strong Generalization with Failure TrajectoriesRuimeng Ye, Zihan Wang, Yang Xiao, Zinan Ling 等ICLR 2026 · 被引用 1 次
- Exploring Semantic-constrained Adversarial Example with Instruction Uncertainty ReductionJin Hu, Jiakai Wang, Linna Jing, Haolin Li 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson 等ICML 2023 · 被引用 908 次
- Cross-Task Generalization via Natural Language Crowdsourcing InstructionsSwaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh HajishirziACL 2022 · 被引用 887 次
相关 Paper
- How to Mitigate Overfitting in Weak-to-strong Generalization?Junhao Shi, Qinyuan Cheng, Zhaoye Fei, Yining Zheng 等ACL 2025 · 被引用 1 次
- Weak to Strong Generalization for Large Language Models with Multi-capabilitiesYucheng Zhou, Jianbing Shen, Yu ChengICLR 2025
- Bayesian WeakS-to-Strong from Text Classification to GenerationZiyun Cui, Ziyang Zhang, Guangzhi Sun, Wen Wu 等ICLR 2025
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker 等ICML 2024 · 被引用 443 次
- A transfer learning framework for weak to strong generalizationSeamus Somerstep, Felipe Maia Polo, Moulinath Banerjee, Yaacov Ritov 等ICLR 2025
