Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models
Yi Ding, Lijun Li, Bing Cao, Jing Shao
摘要
Large Vision-Language Models (VLMs) have achieved remarkable performance across a wide range of tasks. However, their deployment in safety-critical domains poses significant challenges. Existing safety fine-tuning methods, which focus on textual or multimodal content, fall short in addressing challenging cases or disrupt the balance between helpfulness and harmlessness. Our evaluation highlights a safety reasoning gap: these methods lack safety visual reasoning ability, leading to such bottlenecks. To address this limitation and enhance both visual perception and reasoning in safety-critical contexts, we propose a novel dataset that integrates multi-image inputs with safety Chain-of-Thought (CoT) labels as fine-grained reasoning logic to improve model performance. Specifically, we introduce the Multi-Image Safety (MIS) dataset, an instruction-following dataset tailored for multi-image safety scenarios, consisting of training and test splits. Our experiments demonstrate that fine-tuning InternVL2.5-8B with MIS significantly outperforms both powerful open-source models and API-based models in challenging multiimage tasks requiring safety-related visual reasoning. This approach not only delivers exceptional safety performance but also preserves general capabilities without any trade-offs. Specifically, fine-tuning with MIS increases average accuracy by 0.83% across five general benchmarks and reduces the Attack Success Rate (ASR) on multiple safety benchmarks by a large margin. NOTE: This paper contains harmful images & text examples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Backdoor Cleaning without External Guidance in MLLM Fine-tuningXuankun Rong, Wenke Huang, Jian Liang, Jinhe Bi 等NeurIPS 2025 · 被引用 39 次
- Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context InjectionZiqi Miao, Yi Ding, Lijun Li, Jing ShaoEMNLP 2025 · 被引用 21 次
- Sherlock: Self-Correcting Reasoning in Vision-Language ModelsYi Ding, Ruqi ZhangNeurIPS 2025 · 被引用 14 次
- JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language ModelsJiaxin Song, Yixu Wang, Jie Li, Xuan Tong 等NeurIPS 2025 · 被引用 14 次
- SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy OptimizationXuankun Rong, Wenke Huang, Tingfeng Wang, Daiguo Zhou 等CVPR 2026 · 被引用 13 次
它引用的顶会 Paper18
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
相关 Paper
- SURE: Safety Understanding and Reasoning Enhancement for Multimodal Large Language ModelsYuxin Gou, Xiaoning Dong, Qin Li, Shishen Gu 等EMNLP 2025 · 被引用 4 次
- Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related ImagesQishun Yang, Shu Yang, Lijie Hu, Di WangACL 2026 · 被引用 1 次
- Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?Yanbo Wang, Jiyang Guan, Jian Liang, Ran HeCVPR 2025
- Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language ModelsYongshuo Zong, Ondrej Bohdal, Tingyang Yu, Yongxin Yang 等ICML 2024 · 被引用 140 次
- Multimodal Situational SafetyKaiwen Zhou, Chengzhi Liu, Xuandong Zhao, Anderson Compalas 等ICLR 2025
