Lune

USENIX Security2026Top-tier venue

SafeAdapt: Safety Alignment with Adaptive Thinking Allocation for Large Reasoning Models

Jiazheng Song, Junxu Liu, Jian Lou, Jinfei Liu

2026Year

Abstract

Reasoning models have garnered growing importance as their strong Chain-of-Thought (CoT) capabilities unleash exceptional performance on complex tasks. Consequently, an emerging research direction explores "safety thinking" which leverages reasoning models' reasoning capabilities to assess the safety of input prompts and prevent harmful content generation. A prevailing yet under-examined belief in existing studies is that allocating more computational resources to the safety reasoning budget, i.e., rolling out longer safety thinking trajectories, necessarily yields safer behavior. However, this premise has not been systematically analyzed, despite the widespread adoption of reasoning models.

In this paper, we aim to fill this gap by conducting a rigorous safety evaluation to examine the relationship between the safety reasoning budget and the safety of reasoning model generations, particularly under adversarial conditions. Our analysis reveals that safety does not monotonically improve with longer thinking trajectories: both over-thinking on simple attacks and under-thinking on difficult attacks incur excess safety risks, leading to systematic failures under fixed-budget safety policies. Motivated by these key findings, we propose SafeAdapt: Safety Alignment with adaptive thinking Allocation, a novel approach that learns to dynamically adjust reasoning models' safety thinking budget based on prompt difficulty. This enables reasoning models to allocate computational resources adaptively, rather than following a single fixed thinking pattern (e.g., the longer the safer), thereby mitigating excess risks incurred by under-/over-thinking. Experiments across multiple adversarial benchmarks demonstrate that SafeAdapt significantly improves safety performance with a more ideal thinking budget under diverse and challenging attack settings. Qw en 3-0. 6B Qw en 3-1. 7B Qw en 3-4B Qw en 3-8B DS -1 .5 DS -7 B DS -8 B 0.0 0.2 0.4 0.6 0.8 1.0 Proportion of prompts Easy Medium Hard (a) Defense difficulty distribution Qw en 3-0. 6B Qw en 3-1. 7B Qw en 3-4B Qw en 3-8B DS -1 .5 DS -7 B DS -8 B

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Builds on24

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines