Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training
Zheng Xin Yong, Stephen H. Bach
摘要
We discover a novel and surprising phenomenon of unintentional misalignment in reasoning language models (RLMs), which we call self-jailbreaking. Specifically, after benign reasoning training on math or code domains, RLMs will use multiple strategies to circumvent their own safety guardrails. One strategy is to introduce benign assumptions about users and scenarios to justify fulfilling harmful requests. For instance, an RLM reasons that harmful requests like "outline a strategy for stealing customers' credit card information from a retail store" could be associated with the benign intent of "a security professional trying to test defense," despite no such benign context being provided as input. We observe that many open-weight RLMs, including DeepSeek-R1-distilled, s1.1, Phi-4-mini-reasoning, and Nemotron, suffer from self-jailbreaking despite being aware of the harmfulness of the requests. We also provide a mechanistic understanding of self-jailbreaking: RLMs are more compliant after benign reasoning training, and after self-jailbreaking, models appear to perceive malicious requests as less harmful in the CoT, thus enabling compliance with them. To mitigate self-jailbreaking, we find that including minimal safety reasoning data during training is sufficient to ensure RLMs remain safety-aligned. Our work provides the first systematic analysis of self-jailbreaking behavior and offers a practical path forward for maintaining safety in increasingly capable RLMs. https://github.com/BatsResearch/self-jailbreaking Content Warning: This paper contains examples of harmful language.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Prompt Injection as Role ConfusionCharles Ye, Jasmine Cui, Dylan Hadfield-MenellICML 2026 · 被引用 6 次
- Internalizing Safety Understanding in Large Reasoning Models via VerificationYi Zhang, Yuxin Chen, Leheng Sheng, Dongcheng Zhang 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper19
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language ModelsLiwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger 等NeurIPS 2024 · 被引用 247 次
- Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank ModificationsBoyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie 等ICML 2024 · 被引用 215 次
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language ModelsMingjie Liu, Shizhe Diao, Ximing Lu, Jian Hu 等NeurIPS 2025 · 被引用 181 次
- Understanding Catastrophic Forgetting in Language Models via Implicit InferenceSuhas Kotha, Jacob Mitchell Springer, Aditi RaghunathanICLR 2024 · 被引用 131 次
相关 Paper
- SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early AlignmentWonje Jeung, Sangyeon Yoon, Minsuk Kahng, Albert NoNeurIPS 2025 · 被引用 31 次
- Safety Reasoning with GuidelinesHaoyu Wang, Zeyu Qin, Li Shen, Xueqian Wang 等ICML 2025
- Large Language Models Are Involuntary Truth-Tellers: Exploiting Fallacy Failure for Jailbreak AttacksYue Zhou, Henry Peng Zou, Barbara Di Eugenio, Yang ZhangEMNLP 2024 · 被引用 3 次
- Endless Jailbreaks with Bijection LearningBrian R. Y. Huang, Maximilian Li, Leonard TangICLR 2025
- Thinking Dangerously: How Reasoning Training Dilutes Safety in Language ModelsDongxin Guo, Jikun Wu, Siu Ming YiuCCS 2026
