Reasoning-to-Defend: Safety-Aware Reasoning Can Defend Large Language Models from Jailbreaking
Junda Zhu, Lingyong Yan, Shuaiqiang Wang, Dawei Yin, Lei Sha
摘要
Large Reasoning Models (LRMs) have recently demonstrated impressive performances across diverse domains. However, how the safety of Large Language Models (LLMs) benefits from enhanced reasoning capabilities against jailbreak queries remains unexplored. To bridge this gap, in this paper, we propose Reasoningto-Defend (R2D), a novel training paradigm that integrates a safety-aware reasoning mechanism into LLMs' generation process. This enables self-evaluation at each step of the reasoning process, forming safety PIVOT TOKENS as indicators of the safety status of responses. Furthermore, in order to improve the accuracy of predicting PIVOT TOKENS, we propose Contrastive Pivot Optimization (CPO), which enhances the model's perception of the safety status of given dialogues. LLMs dynamically adjust their response strategies during reasoning, significantly enhancing their safety capabilities defending jailbreak attacks. Extensive experiments demonstrate that R2D effectively mitigates various attacks and improves overall safety, while maintaining the original performances. This highlights the substantial potential of safety-aware reasoning in improving robustness of LRMs and LLMs against various jailbreaks. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Reasoning as an Adaptive Defense for SafetyTaeyoun Kim, Fahim Tajwar, Aditi Raghunathan, Aviral KumarNeurIPS 2025 · 被引用 24 次
- AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement LearningYi Zhang, An Zhang, XiuYu Zhang, Leheng Sheng 等ICLR 2026 · 被引用 15 次
- Towards Safe Reasoning in Large Reasoning Models via Corrective InterventionYichi Zhang, Yue Ding, Jingwen Yang, Tianwei Luo 等ICLR 2026 · 被引用 13 次
- Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning TrainingZheng Xin Yong, Stephen H. BachICLR 2026 · 被引用 10 次
- CARE: Decoding-Time Safety Alignment via Rollback and Introspection InterventionXiaomeng Hu, Fei Huang, Chenhan Yuan, Junyang Lin 等NeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper26
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
相关 Paper
- ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha MomentsYuquan Wang, Mi Zhang, Yining Wang, Geng Hong 等ACL 2026 · 被引用 2 次
- SafeKey: Amplifying Aha-Moment Insights for Safety ReasoningKaiwen Zhou, Xuandong Zhao, Jayanth Srinivasa, Gaowen Liu 等EMNLP 2025
- ARMOR: Aligning Secure and Safe Large Language Models via Meticulous ReasoningZhengyue Zhao, Yingzi Ma, Somesh Jha, Marco Pavone 等ICLR 2026 · 被引用 8 次
- Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language ModelsGuangyu Yang, Jinghong Chen, Jingbiao Mei, Weizhe Lin 等ACL 2026 · 被引用 1 次
- Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time AlignmentSoumya Suvra Ghosal, Souradip Chakraborty, Vaibhav Singh, Tianrui Guan 等CVPR 2025
