Safety Alignment Can Be Not Superficial With Explicit Safety Signals
Jianwei Li, Jung-Eun Kim
Abstract
Recent studies on the safety alignment of large language models (LLMs) have revealed that existing approaches often operate superficially, leaving models vulnerable to various adversarial attacks. Despite their significance, these studies generally fail to offer actionable solutions beyond data augmentation for achieving more robust safety mechanisms. This paper identifies a fundamental cause of this superficiality: existing alignment approaches often presume that models can implicitly learn a safety-related reasoning task during the alignment process, enabling them to refuse harmful requests. However, the learned safety signals are often diluted by other competing objectives, leading models to struggle with drawing a firm safety-conscious decision boundary when confronted with adversarial attacks. Based on this observation, by explicitly introducing a safety-related binary classification task and integrating its signals with our attention and decoding strategies, we eliminate this ambiguity and allow models to respond more responsibly to malicious queries. We emphasize that, with less than 0.2x overhead cost, our approach enables LLMs to assess the safety of both the query and the previously generated tokens at each necessary generating step. Extensive experiments demonstrate that our method significantly improves the resilience of LLMs against various adversarial attacks, offering a promising pathway toward more robust generative AI systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b2416d07-e971-4fb7-a2f0-cd0e58c2c701Cited by top-tier papers5
- Superficial Safety Alignment HypothesisJianwei Li, Jung-Eun KimICLR 2026 · 11 citations
- Purifying Generative LLMs from Backdoors without Prior Knowledge or Clean ReferenceJianwei Li, Jung-Eun KimICLR 2026 · 8 citations
- Safety Alignment of Large Language Models via Contrasting Safe and Harmful DistributionsXiaoyun Zhang, Zhengyue Zhao, Wenxuan Shi, Kaidi Xu et al.AAAI 2026 · 4 citations
- Alignment-Weighted DPO: A principled reasoning approach to improve safety alignmentMengxuan Hu, Vivek V. Datla, Anoop Kumar, Zihan Guan et al.ICLR 2026 · 3 citations
- New Wide-Net-Casting Jailbreak Attacks Risk Large ModelsQiuchi Xiang, Haoxuan Qu, Hossein Rahmani, Jun LiuICML 2026
Builds on26
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
Related papers
- Resolving the Security-Auditability Dilemma with Auditable Latent Chain-of-Thought AlignmentGuan Wang, Biyu Zhou, Xuehai Tang, Jizhong Han et al.ACL 2026
- Safety Alignment Should be Made More Than Just a Few Tokens DeepXiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma et al.ICLR 2025
- AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement LearningYi Zhang, An Zhang, XiuYu Zhang, Leheng Sheng et al.ICLR 2026 · 15 citations
- Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-DepthJiawei Zhang, Andrew Estornell, David D. Baek, Bo Li et al.ICLR 2026 · 3 citations
- Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-CheckChentao Cao, Xiaojun Xu, Bo Han, Hang LiICLR 2026 · 5 citations
