Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization
Wenkai Yang, Shiqi Shen, Guangyao Shen, Wei Yao, Yong Liu, Gong Zhi, Yankai Lin, Ji-Rong Wen
Abstract
Superalignment, where humans act as weak supervisors for superhuman models, has become a crucial problem with the rapid development of Large Language Models (LLMs). Recent work has preliminarily studied this problem by using weak models to supervise strong models, and discovered that weakly supervised strong students can consistently outperform weak teachers towards the alignment target, leading to a weak-to-strong generalization phenomenon. However, we are concerned that behind such a promising phenomenon, whether there exists an issue of weak-to-strong deception, where strong models deceive weak models by exhibiting well-aligned in areas known to weak models but producing misaligned behaviors in cases weak models do not know. We take an initial step towards exploring this security issue in a specific but realistic multi-objective alignment case, where there may be some alignment targets conflicting with each other (e.g., helpfulness v.s. harmlessness). We aim to explore whether, in such cases, strong models might deliberately make mistakes in areas known to them but unknown to weak models within one alignment dimension, in exchange for a higher reward in another dimension. Through extensive experiments in both the reward modeling and preference optimization scenarios, we find: (1) The weak-to-strong deception phenomenon exists across all settings. (2) The deception intensifies as the capability gap between weak and strong models increases. (3) Bootstrapping with an intermediate model can mitigate the deception to some extent, though its effectiveness remains limited. Our work highlights the urgent need to pay more attention to the true reliability of superalignment. 1 * Corresponding Author 1 Code is available at https://github.com/RUCBM/weak-to-strong-deception . Human Superhuman Model (a) Superalignment supervise Weak Model Strong Model supervise (b) Analogous Problem Weak Model Strong Model Strong Model W ea k K n o w n W ea k U n k n o w n St ro ng K no w n St ro ng U nk no w n W ea k K n o w n W ea k U n k n o w n St ro ng K no w n St ro ng U nk no w n s u p e r v is e s u p e r v is e (c) Weak-to-Strong Generalization (d) Weak-to-Strong Deception generate weak supervision data W ea k K n o w n W ea k U n k n o w n St ro ng K no w n St ro ng U nk no w n another conflicting alignment target Generalizing well to areas unknow to the weak model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7577223d-42ae-46b0-b2eb-90ea45f3f7e3Cited by top-tier papers10
- Weak-to-Strong Search: Align Large Language Models via Searching over Small Language ModelsZhanhui Zhou, Zhixuan Liu, Jie Liu, Zhichen Dong et al.NeurIPS 2024 · 50 citations
- Contrastive Weak-to-Strong GeneralizationHoucheng Jiang, Junfeng Fang, Jiaxin Wu, Tianyu Zhang et al.ICML 2026 · 2 citations
- Does Weak-to-strong Generalization Happen under Spurious Correlations?Chenruo Liu, Yijun Dong, Qi LeiICLR 2026 · 1 citation
- Weak-to-Strong Generalization with Failure TrajectoriesRuimeng Ye, Zihan Wang, Yang Xiao, Zinan Ling et al.ICLR 2026 · 1 citation
- Exploring Semantic-constrained Adversarial Example with Instruction Uncertainty ReductionJin Hu, Jiakai Wang, Linna Jing, Haolin Li et al.NeurIPS 2025 · 1 citation
Builds on14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson et al.ICML 2023 · 908 citations
- Cross-Task Generalization via Natural Language Crowdsourcing InstructionsSwaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh HajishirziACL 2022 · 887 citations
Related papers
- How to Mitigate Overfitting in Weak-to-strong Generalization?Junhao Shi, Qinyuan Cheng, Zhaoye Fei, Yining Zheng et al.ACL 2025 · 1 citation
- Weak to Strong Generalization for Large Language Models with Multi-capabilitiesYucheng Zhou, Jianbing Shen, Yu ChengICLR 2025
- Bayesian WeakS-to-Strong from Text Classification to GenerationZiyun Cui, Ziyang Zhang, Guangzhi Sun, Wen Wu et al.ICLR 2025
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker et al.ICML 2024 · 443 citations
- A transfer learning framework for weak to strong generalizationSeamus Somerstep, Felipe Maia Polo, Moulinath Banerjee, Yaacov Ritov et al.ICLR 2025
