USENIX Security2026Top-tier venue
SafeAdapt: Safety Alignment with Adaptive Thinking Allocation for Large Reasoning Models
Jiazheng Song, Junxu Liu, Jian Lou, Jinfei Liu
Abstract
Reasoning models have garnered growing importance as their strong Chain-of-Thought (CoT) capabilities unleash exceptional performance on complex tasks. Consequently, an emerging research direction explores "safety thinking" which leverages reasoning models' reasoning capabilities to assess the safety of input prompts and prevent harmful content generation. A prevailing yet under-examined belief in existing studies is that allocating more computational resources to the safety reasoning budget, i.e., rolling out longer safety thinking trajectories, necessarily yields safer behavior. However, this premise has not been systematically analyzed, despite the widespread adoption of reasoning models.
In this paper, we aim to fill this gap by conducting a rigorous safety evaluation to examine the relationship between the safety reasoning budget and the safety of reasoning model generations, particularly under adversarial conditions. Our analysis reveals that safety does not monotonically improve with longer thinking trajectories: both over-thinking on simple attacks and under-thinking on difficult attacks incur excess safety risks, leading to systematic failures under fixed-budget safety policies. Motivated by these key findings, we propose SafeAdapt: Safety Alignment with adaptive thinking Allocation, a novel approach that learns to dynamically adjust reasoning models' safety thinking budget based on prompt difficulty. This enables reasoning models to allocate computational resources adaptively, rather than following a single fixed thinking pattern (e.g., the longer the safer), thereby mitigating excess risks incurred by under-/over-thinking. Experiments across multiple adversarial benchmarks demonstrate that SafeAdapt significantly improves safety performance with a more ideal thinking budget under diverse and challenging attack settings. Qw en 3-0. 6B Qw en 3-1. 7B Qw en 3-4B Qw en 3-8B DS -1 .5 DS -7 B DS -8 B 0.0 0.2 0.4 0.6 0.8 1.0 Proportion of prompts Easy Medium Hard (a) Defense difficulty distribution Qw en 3-0. 6B Qw en 3-1. 7B Qw en 3-4B Qw en 3-8B DS -1 .5 DS -7 B DS -8 B
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on24
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang et al.NeurIPS 2025 · 1,109 citations
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language ModelsXiaogeng Liu, Nan Xu, Muhao Chen, Chaowei XiaoICLR 2024 · 722 citations
Related papers
- Finding and Reactivating Post-Trained LLMs' Hidden Safety MechanismsMingjie Li, Wai Man Si, Michael Backes, Yang Zhang et al.NeurIPS 2025 · 4 citations
- Teach to Reason Safely: Policy-Guided Safety Tuning for MLRMsJingyu Zhang, Kun Yang, Ming Wen, Zhuoer Xu et al.ICLR 2026
- AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning ModelsZihao Zhu, Xinyu Wu, Gehan Hu, Siwei Lyu et al.ICLR 2026 · 6 citations
- Safety Alignment Can Be Not Superficial With Explicit Safety SignalsJianwei Li, Jung-Eun KimICML 2025
- Towards Safe Reasoning in Large Reasoning Models via Corrective InterventionYichi Zhang, Yue Ding, Jingwen Yang, Tianwei Luo et al.ICLR 2026 · 13 citations
