Teach to Reason Safely: Policy-Guided Safety Tuning for MLRMs
Jingyu Zhang, Kun Yang, Ming Wen, Zhuoer Xu, Zeyang Sha, Shiwen Cui, Zhaohui Yang
摘要
Multimodal Large Reasoning Models (MLRMs) have exhibited remarkable capabilities in complex multimodal tasks. However, our findings reveal a critical trade-off: reasoning-based models are more prone to generating harmful content, leading to degradation in safety performance. This paper presents a large-scale analysis of this safety–reasoning trade-off, identifying two main mechanisms of safety degradation: (i) visual attention drift, which reduces the model’s reliance on visual grounding and thereby exacerbates overlooked risks in cross-modal interactions; (ii) unsafe reasoning patterns, including flawed reasoning initiation and chain-of-thought safety attenuation, which compromise the model’s safety awareness. To mitigate these issues, we propose Policy-guided Safety Tuning (PST), a two-stage alignment framework. It first employs Policy-Guided Supervised Fine-Tuning to integrate explicit safety policies into the reasoning process, establishing a structured and interpretable foundation for safe decision-making. Then, PST applies Safety Reasoning Preference Optimization to encourage the model to generate safe, helpful, and informative responses while reducing oversensitive and homogeneous characteristics. Extensive experiments demonstrate that PST significantly reduces harmful outputs across multiple multimodal safety benchmarks, while maintaining competitive performance on general tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
- More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning ModelsZhongxing Xu, Chengzhi Liu, Qingyue Wei, Juncheng Wu 等NeurIPS 2025 · 被引用 103 次
- RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-Grained Correctional Human FeedbackTianyu Yu, Yuan Yao, Haoye Zhang, Taiwen He 等CVPR 2024 · 被引用 72 次
- LlaVA-CoT: Let Vision Language Models Reason Step-By-StepGuowei Xu, Peng Jin, Ziang Wu, Hao Li 等ICCV 2025 · 被引用 37 次
相关 Paper
- Pragma-VL: Towards a Pragmatic Arbitration of Safety and Helpfulness in MLLMsMing Wen, Kun Yang, Xin Chen, Jingyu Zhang 等ICLR 2026 · 被引用 4 次
- Think in Safety: Unveiling and Mitigating Safety Alignment Collapse in Multimodal Large Reasoning ModelXinyue Lou, You Li, Jinan Xu, Xiangyu Shi 等EMNLP 2025
- Mitigating Safety Context Amnesia in Multimodal Reasoning Models via Intent-Guided Safety ReasoningXiyao Dong, Guangsheng Cheng, YiLong Chen, Xiaojin Zhang 等ACL 2026
- RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguardsJingnan Zheng, Xiangtian Ji, Yijun Lu, Chenhang Cui 等NeurIPS 2025 · 被引用 17 次
- Reasoning Structure Matters for Safety Alignment of Reasoning ModelsYeonjun In, Wonjoong Kim, Sangwu Park, Chanyoung ParkACL 2026
