Reasoning as an Adaptive Defense for Safety
Taeyoun Kim, Fahim Tajwar, Aditi Raghunathan, Aviral Kumar
Abstract
Reasoning methods that adaptively allocate test-time compute have advanced LLM performance on easy to verify domains such as math and code. In this work, we study how to utilize this approach to train models that exhibit a degree of robustness to safety vulnerabilities, and show that doing so can provide benefits. We build a recipe called (Training Adaptive Reasoners for Safety), a reinforcement learning (RL) approach that trains models to reason about safety using chain-of-thought traces and a reward signal that balances safety with task completion. To build TARS, we identify three critical design choices: (1) a ``lightweight''warmstart SFT stage, (2) a mix of harmful, harmless, and ambiguous prompts to prevent shortcut behaviors such as too many refusals, and (3) a reward function to prevent degeneration of reasoning capabilities during training. Models trained with TARS exhibit adaptive behaviors by spending more compute on ambiguous queries, leading to better safety-refusal trade-offs. They also internally learn to better distinguish between safe and unsafe prompts and attain greater robustness to both white-box (e.g., GCG) and black-box attacks (e.g., PAIR). Overall, our work provides an effective, open recipe for training LLMs against jailbreaks and harmful requests by reasoning per prompt.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal JailbreaksHoagy Cunningham, Jerry Wei, Zihan Wang, Andrew Persic et al.ICLR 2026 · 39 citations
- Towards Safe Reasoning in Large Reasoning Models via Corrective InterventionYichi Zhang, Yue Ding, Jingwen Yang, Tianwei Luo et al.ICLR 2026 · 13 citations
- Mitigating the Safety–Utility Trade-off in LLM Alignment via Adaptive Safe Context LearningYanbo Wang, Minzheng Wang, Jian Liang, Lu Wang et al.ICML 2026 · 3 citations
- AutoRAN: Automated Hijacking of Safety Reasoning in Large Reasoning ModelsJiacheng Liang, Tanqiu Jiang, Yuhui Wang, Rongyi Zhu et al.ACL 2026 · 3 citations
- Submodular Optimization for Minimal Augmentation in Robust Language Model AlignmentCHING-CHIA KAO, Chia-Mu Yu, Chun-Shien Lu, Chu-song ChenICML 2026
Builds on28
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
Related papers
- Safety Reasoning with GuidelinesHaoyu Wang, Zeyu Qin, Li Shen, Xueqian Wang et al.ICML 2025
- Endless Jailbreaks with Bijection LearningBrian R. Y. Huang, Maximilian Li, Leonard TangICLR 2025
- The Alignment Waltz: Jointly Training Agents to Collaborate for SafetyJingyu Zhang, Haozhu Wang, Eric Michael Smith, Sid Wang et al.ICLR 2026 · 11 citations
- Alignment-Weighted DPO: A principled reasoning approach to improve safety alignmentMengxuan Hu, Vivek V. Datla, Anoop Kumar, Zihan Guan et al.ICLR 2026 · 3 citations
- Adversarial Reasoning at Jailbreaking TimeMahdi Sabbaghi, Paul Kassianik, George J. Pappas, Amin Karbasi et al.ICML 2025
