RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards
Jingnan Zheng, Xiangtian Ji, Yijun Lu, Chenhang Cui, Weixiang Zhao, Gelei Deng, Zhenkai Liang, An Zhang, Tat-Seng Chua
Abstract
Large Language Models (LLMs) continue to exhibit vulnerabilities despite deliberate safety alignment efforts, posing significant risks to users and society. To safeguard against the risk of policy-violating content, system-level moderation via external guard models-designed to monitor LLM inputs and outputs and block potentially harmful content-has emerged as a prevalent mitigation strategy. Existing approaches of training guard models rely heavily on extensive human curated datasets and struggle with out-of-distribution threats, such as emerging harmful categories or jailbreak attacks. To address these limitations, we propose RSafe, an adaptive reasoning-based safeguard that conducts guided safety reasoning to provide robust protection within the scope of specified safety policies. RSafe operates in two stages: (1) guided reasoning, where it analyzes safety risks of input content through policy-guided step-by-step reasoning, and (2) reinforced alignment, where rule-based RL optimizes its reasoning paths to align with accurate safety prediction. This two-stage training paradigm enables RSafe to internalize safety principles to generalize safety protection capability over unseen or adversarial safety violation scenarios. During inference, RSafe accepts user-specified safety policies to provide enhanced safeguards tailored to specific safety requirements. Experiments demonstrate that RSafe matches state-of-the-art guard models using limited amount of public data in both prompt-and response-level harmfulness detection, while achieving superior out-of-distribution generalization on both emerging harmful category and jailbreak attacks. Furthermore, RSafe provides human-readable explanations for its safety judgments for better interpretability. RSafe offers a robust, adaptive, and interpretable solution for LLM safety moderation, advancing the development of reliable safeguards in dynamic real-world environments. Our code is available at https://github.com/SophieZheng998/RSafe.git
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47813581-50c0-43cd-84db-c45b21c7bed1Cited by top-tier papers7
- AlphaSteer: Learning Refusal Steering with Principled Null-Space ConstraintLeheng Sheng, Changshuo Shen, Weixiang Zhao, Junfeng Fang et al.ICLR 2026 · 52 citations
- Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language ModelsChenhang Cui, Gelei Deng, An Zhang, Jingnan Zheng et al.NeurIPS 2025 · 10 citations
- Alignment-Weighted DPO: A principled reasoning approach to improve safety alignmentMengxuan Hu, Vivek V. Datla, Anoop Kumar, Zihan Guan et al.ICLR 2026 · 3 citations
- Internalizing Safety Understanding in Large Reasoning Models via VerificationYi Zhang, Yuxin Chen, Leheng Sheng, Dongcheng Zhang et al.ICML 2026 · 2 citations
- FlexGuard: Continuous Risk Scoring for Strictness-Adaptive LLM Content ModerationZhihao Ding, Jinming Li, Ze Lu, Jieming ShiACL 2026 · 2 citations
Builds on21
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?Zhiqi Chen, Rui Lu, Andrew Zhao, Zhaokai Wang et al.NeurIPS 2025 · 1,109 citations
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
Related papers
- MrGuard: A Multilingual Reasoning Guardrail for Universal LLM SafetyYahan Yang, Soham Dan, Shuo Li, Dan Roth et al.EMNLP 2025
- ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha MomentsYuquan Wang, Mi Zhang, Yining Wang, Geng Hong et al.ACL 2026 · 2 citations
- R2-Guard: Robust Reasoning Enabled LLM Guardrail via Knowledge-Enhanced Logical ReasoningMintong Kang, Bo LiICLR 2025
- SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early AlignmentWonje Jeung, Sangyeon Yoon, Minsuk Kahng, Albert NoNeurIPS 2025 · 31 citations
- ARMOR: Aligning Secure and Safe Large Language Models via Meticulous ReasoningZhengyue Zhao, Yingzi Ma, Somesh Jha, Marco Pavone et al.ICLR 2026 · 8 citations
