SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Models
Yongting Zhang, Lu Chen, Guodong Zheng, Yifeng Gao, Rui Zheng, Jinlan Fu, Zhenfei Yin, Senjie Jin, Yu Qiao, Xuanjing Huang, Feng Zhao, Tao Gui, Jing Shao
Abstract
The emergence of Vision Language Models (VLMs) has brought unprecedented advances in understanding multimodal information. The combination of textual and visual semantics in VLMs is highly complex and diverse, making the safety alignment of these models challenging. Furthermore, due to the limited study on the safety alignment of VLMs, there is a lack of large-scale, high-quality datasets. To address these limitations, we propose a Safety Preference Alignment dataset for Vision Language Models named SPA-VL. In terms of breadth, SPA-VL covers 6 harmfulness domains, 13 categories, and 53 subcategories, and contains 100,788 samples of the quadruple (question, image, chosen response, rejected response). In terms of depth, the responses are collected from 12 opensource (e.g., QwenVL) and closed-source (e.g., Gemini) VLMs to ensure diversity. The construction of preference data is fully automated, and the experimental results indicate that models trained with alignment techniques on the SPA-VL dataset exhibit substantial improvements in harmlessness and helpfulness while maintaining core capabilities. SPA-VL, as a large-scale, high-quality, and diverse dataset, represents a significant milestone in ensuring that VLMs achieve both harmlessness and helpfulness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b27a1c6b-12ee-42b8-b816-d9265dfd0899Cited by top-tier papers18
- Understanding and Rectifying Safety Perception Distortion in VLMsXiaohan Zou, Jian Kang, George Kesidis, Lu LinNeurIPS 2025 · 20 citations
- SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy OptimizationXuankun Rong, Wenke Huang, Tingfeng Wang, Daiguo Zhou et al.CVPR 2026 · 13 citations
- Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive ScoringPeichun Hua, Hao Li, Shanghao Shi, Zhiyuan Yu et al.ACL 2026 · 8 citations
- Differentially Private Federated Low Rank Adaptation Beyond Fixed-MatrixMing Wen, Jiaqi Zhu, Yuedong Xu, Yipeng Zhou et al.NeurIPS 2025 · 6 citations
- SARSteer: Safeguarding Large Audio Language Models via Safe-Ablated Refusal SteeringWeilin Lin, Jianze Li, Hui Xiong, Li LiuICML 2026 · 6 citations
Builds on19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human PreferenceJiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen et al.ACL 2025
- Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?Yanbo Wang, Jiyang Guan, Jian Liang, Ran HeCVPR 2025
- Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related ImagesQishun Yang, Shu Yang, Lijie Hu, Di WangACL 2026 · 1 citation
- ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference TimeYi Ding, Bolian Li, Ruqi ZhangICLR 2025 · 2 citations
- Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human FeedbackJiaming Ji, Xinyu Chen, Rui Pan, Han Zhu et al.NeurIPS 2025 · 28 citations
