Improving LLM Safety Alignment with Dual-Objective Optimization
Xuandong Zhao, Will Cai, Tianneng Shi, David Huang, Licong Lin, Song Mei, Dawn Song
摘要
Existing training-time safety alignment techniques for large language models (LLMs) remain vulnerable to jailbreak attacks. Direct preference optimization (DPO), a widely deployed alignment method, exhibits limitations in both experimental and theoretical contexts as its loss function proves suboptimal for refusal learning. Through gradient-based analysis, we identify these shortcomings and propose an improved safety alignment that disentangles DPO objectives into two components: (1) robust refusal training, which encourages refusal even when partial unsafe generations are produced, and (2) targeted unlearning of harmful knowledge. This approach significantly increases LLM robustness against a wide range of jailbreak attacks, including prefilling, suffix, and multi-turn attacks across both in-distribution and out-of-distribution scenarios. Furthermore, we introduce a method to emphasize critical refusal tokens by incorporating a reward-based tokenlevel weighting mechanism for refusal learning, which further improves the robustness against adversarial exploits. Our research also suggests that robustness to jailbreak attacks is correlated with token distribution shifts in the training process and internal representations of refusal and harmful tokens, offering valuable directions for future research in LLM safety alignment. The code is available at https://github.com/ wicai24/DOOR-Alignment .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Mitigating the Safety Alignment Tax with Null-Space Constrained Policy OptimizationYifan Niu, Han Xiao, Dongyi Liu, Nuo Chen 等ICLR 2026 · 被引用 13 次
- Offline RL by Reward-Weighted Fine-Tuning for Conversation OptimizationSubhojyoti Mukherjee, Viet Dac Lai, Raghavendra Addanki, Ryan Rossi 等NeurIPS 2025 · 被引用 12 次
- Model Unlearning via Sparse Autoencoder Subspace Guided ProjectionsXu Wang, Zihao Li, Benyou Wang, Yan Hu 等EMNLP 2025 · 被引用 9 次
- Toward Safer Diffusion Language Models: Discovery and Mitigation of Priming VulnerabilityShojiro Yamabe, Jun SakumaICLR 2026 · 被引用 9 次
- When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured OutputShuoming Zhang, Jiacheng Zhao, Hanyuan Dong, Ruiyuan Xu 等CCS 2026 · 被引用 6 次
它引用的顶会 Paper16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 被引用 722 次
- Improving Alignment and Robustness with Circuit BreakersAndy Zou, Long Phan, Justin Wang, Derek Duenas 等NeurIPS 2024 · 被引用 362 次
相关 Paper
- Alignment-Weighted DPO: A principled reasoning approach to improve safety alignmentMengxuan Hu, Vivek V. Datla, Anoop Kumar, Zihan Guan 等ICLR 2026 · 被引用 3 次
- TAO-Attack: Toward Advanced Optimization-Based Jailbreak Attacks for Large Language ModelsZhi Xu, Jiaqi Li, Xiaotong Zhang, Hong Yu 等ICLR 2026 · 被引用 2 次
- Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template RegionChak Tou Leong, Qingyu Yin, Jian Wang, Wenjie LiACL 2025
- SHARP: Self-adaptive Harmful Category-aware Prompt Generation for Black-box JailbreakingYingjie Xue, Xingyou Xia, Jun Zhang, Yunbo Cao 等ACL 2026
- DualEdit: Mitigating Safety Fallback in LLM Backdoor Editing via Affirmation-Refusal RegulationHoucheng Jiang, Zetong Zhao, Junfeng Fang, Haokai Ma 等ICLR 2026 · 被引用 2 次
