ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time
Yi Ding, Bolian Li, Ruqi Zhang
摘要
Vision Language Models (VLMs) have become essential backbones for multimodal intelligence, yet significant safety challenges limit their real-world application. While textual inputs are often effectively safeguarded, adversarial visual inputs can easily bypass VLM defense mechanisms. Existing defense methods are either resource-intensive, requiring substantial data and compute, or fail to simultaneously ensure safety and usefulness in responses. To address these limitations, we propose a novel two-phase inference-time alignment framework, Evaluating Then Aligning (ETA): 1) Evaluating input visual contents and output responses to establish a robust safety awareness in multimodal settings, and 2) Aligning unsafe behaviors at both shallow and deep levels by conditioning the VLMs' generative distribution with an interference prefix and performing sentence-level best-of-N to search the most harmless and helpful generation paths. Extensive experiments show that ETA outperforms baseline methods in terms of harmlessness, helpfulness, and efficiency, reducing the unsafe rate by 87.5% in cross-modality attacks and achieving 96.6% win-ties in GPT-4 helpfulness evaluation. The code is publicly available at https://github.com/DripNowhy/ETA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- GuardReasoner-VL: Safeguarding VLMs via Reinforced ReasoningYue Liu, Shengfang Zhai, Mingzhe Du, Yulin Chen 等NeurIPS 2025 · 被引用 40 次
- The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and DefenseYangyang Guo, Fangkai Jiao, Liqiang Nie, Mohan KankanhalliNeurIPS 2025 · 被引用 24 次
- Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language ModelsYi Ding, Lijun Li, Bing Cao, Jing ShaoICLR 2026 · 被引用 21 次
- Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context InjectionZiqi Miao, Yi Ding, Lijun Li, Jing ShaoEMNLP 2025 · 被引用 21 次
- Understanding and Rectifying Safety Perception Distortion in VLMsXiaohan Zou, Jian Kang, George Kesidis, Lu LinNeurIPS 2025 · 被引用 20 次
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu 等NeurIPS 2022 · 被引用 2,727 次
相关 Paper
- Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?Yanbo Wang, Jiyang Guan, Jian Liang, Ran HeCVPR 2025
- VLSBench: Unveiling Visual Leakage in Multimodal SafetyXuhao Hu, Dongrui Liu, Hao Li, Xuanjing Huang 等ACL 2025
- Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related ImagesQishun Yang, Shu Yang, Lijie Hu, Di WangACL 2026 · 被引用 1 次
- Harnessing Hyperbolic Geometry for Harmful Prompt Detection and SanitizationIgor Maljkovic, Maria Rosaria Briglia, Iacopo Masi, Antonio Emanuele Cinà 等ICLR 2026 · 被引用 2 次
- GuardAlign: Test-time Safety Alignment in Multimodal Large Language ModelsXingyu Zhu, Beier Zhu, Junfeng Fang, Shuo Wang 等ICLR 2026 · 被引用 2 次
