Reasoning Structure Matters for Safety Alignment of Reasoning Models
Yeonjun In, Wonjoong Kim, Sangwu Park, Chanyoung Park
Abstract
Large reasoning models (LRMs) achieve strong performance on complex reasoning tasks but often generate harmful responses to malicious user queries. This paper investigates the underlying cause of these safety risks and shows that the issue lies in the reasoning structure itself. Based on this insight, we claim that effective safety alignment can be achieved by altering the reasoning structure. We propose ALTTRAIN, a simple yet effective post-training method that explicitly alters the reasoning structure of LRMs. ALTTRAIN is both practical and generalizable, requiring no complex reinforcement learning (RL) training or reward design-only supervised fine-tuning (SFT) with a lightweight 1K training examples. Experiments across LRM backbones and model sizes demonstrate strong safety alignment, along with robust generalization across reasoning, QA, summarization, and multilingual setting. Our code are available at https://github.com/yeonjunin/R1- Alt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8ea8872-270c-4be5-ba08-f0615f8a62e5Builds on9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 2,230 citations
- ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERTOmar Khattab, Matei ZahariaSIGIR 2020 · 1,246 citations
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language ModelsLiwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger et al.NeurIPS 2024 · 247 citations
Related papers
- Finding and Reactivating Post-Trained LLMs' Hidden Safety MechanismsMingjie Li, Wai Man Si, Michael Backes, Yang Zhang et al.NeurIPS 2025 · 4 citations
- SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early AlignmentWonje Jeung, Sangyeon Yoon, Minsuk Kahng, Albert NoNeurIPS 2025 · 31 citations
- AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement LearningYi Zhang, An Zhang, XiuYu Zhang, Leheng Sheng et al.ICLR 2026 · 15 citations
- How Should We Enhance the Safety of Large Reasoning Models: An Empirical StudyZhexin Zhang, Xian Qi Loye, Victor Shea-Jay Huang, Junxiao Yang et al.ACL 2026 · 20 citations
- SafeKey: Amplifying Aha-Moment Insights for Safety ReasoningKaiwen Zhou, Xuandong Zhao, Jayanth Srinivasa, Gaowen Liu et al.EMNLP 2025
