Lune

USENIX Security2026顶会

SafeAdapt: Safety Alignment with Adaptive Thinking Allocation for Large Reasoning Models

Jiazheng Song, Junxu Liu, Jian Lou, Jinfei Liu

出版方
2026年份

摘要

Reasoning models have garnered growing importance as their strong Chain-of-Thought (CoT) capabilities unleash exceptional performance on complex tasks. Consequently, an emerging research direction explores "safety thinking" which leverages reasoning models' reasoning capabilities to assess the safety of input prompts and prevent harmful content generation. A prevailing yet under-examined belief in existing studies is that allocating more computational resources to the safety reasoning budget, i.e., rolling out longer safety thinking trajectories, necessarily yields safer behavior. However, this premise has not been systematically analyzed, despite the widespread adoption of reasoning models.

In this paper, we aim to fill this gap by conducting a rigorous safety evaluation to examine the relationship between the safety reasoning budget and the safety of reasoning model generations, particularly under adversarial conditions. Our analysis reveals that safety does not monotonically improve with longer thinking trajectories: both over-thinking on simple attacks and under-thinking on difficult attacks incur excess safety risks, leading to systematic failures under fixed-budget safety policies. Motivated by these key findings, we propose SafeAdapt: Safety Alignment with adaptive thinking Allocation, a novel approach that learns to dynamically adjust reasoning models' safety thinking budget based on prompt difficulty. This enables reasoning models to allocate computational resources adaptively, rather than following a single fixed thinking pattern (e.g., the longer the safer), thereby mitigating excess risks incurred by under-/over-thinking. Experiments across multiple adversarial benchmarks demonstrate that SafeAdapt significantly improves safety performance with a more ideal thinking budget under diverse and challenging attack settings. Qw en 3-0. 6B Qw en 3-1. 7B Qw en 3-4B Qw en 3-8B DS -1 .5 DS -7 B DS -8 B 0.0 0.2 0.4 0.6 0.8 1.0 Proportion of prompts Easy Medium Hard (a) Defense difficulty distribution Qw en 3-0. 6B Qw en 3-1. 7B Qw en 3-4B Qw en 3-8B DS -1 .5 DS -7 B DS -8 B

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper24

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖