Lune

ICLR2026顶会

Teach to Reason Safely: Policy-Guided Safety Tuning for MLRMs

Jingyu Zhang, Kun Yang, Ming Wen, Zhuoer Xu, Zeyang Sha, Shiwen Cui, Zhaohui Yang

出版方
2026年份

摘要

Multimodal Large Reasoning Models (MLRMs) have exhibited remarkable capabilities in complex multimodal tasks. However, our findings reveal a critical trade-off: reasoning-based models are more prone to generating harmful content, leading to degradation in safety performance. This paper presents a large-scale analysis of this safety–reasoning trade-off, identifying two main mechanisms of safety degradation: (i) visual attention drift, which reduces the model’s reliance on visual grounding and thereby exacerbates overlooked risks in cross-modal interactions; (ii) unsafe reasoning patterns, including flawed reasoning initiation and chain-of-thought safety attenuation, which compromise the model’s safety awareness. To mitigate these issues, we propose Policy-guided Safety Tuning (PST), a two-stage alignment framework. It first employs Policy-Guided Supervised Fine-Tuning to integrate explicit safety policies into the reasoning process, establishing a structured and interpretable foundation for safe decision-making. Then, PST applies Safety Reasoning Preference Optimization to encourage the model to generate safe, helpful, and informative responses while reducing oversensitive and homogeneous characteristics. Extensive experiments demonstrate that PST significantly reduces harmful outputs across multiple multimodal safety benchmarks, while maintaining competitive performance on general tasks.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 341be4e4-518c-4468-8f1e-e7f10ec99bd5

它引用的顶会 Paper10

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖