TRUST-VLM: Thorough Red-Teaming for Uncovering Safety Threats in Vision-Language Models
Kangjie Chen, Muyang Li, Guanlin Li, Shudong Zhang, Shangwei Guo, Tianwei Zhang
摘要
Vision-Language Models (VLMs) have become a cornerstone in multi-modal artificial intelligence, enabling seamless integration of visual and textual information for tasks such as image captioning, visual question answering, and cross-modal retrieval. Despite their impressive capabilities, these models often exhibit inherent vulnerabilities that can lead to safety failures in critical applications. Red-teaming is an important approach to identify and test system's vulnerabilities, but how to conduct red-teaming for contemporary VLMs is an unexplored area. In this paper, we propose a novel multi-modal red-teaming approach, TRUST-VLM, to enhance both the attack success rate and the diversity of successful test cases for VLMs. Specifically, TRUST-VLM is built upon the incontext learning to adversarially test a VLM on both image and text inputs. Furthermore, we involve feedback from the target VLM to improve the efficiency of test case generation. Extensive experiments show that TRUST-VLM not only outperforms traditional red-teaming techniques in generating diverse and effective adversarial cases but also provides actionable insights for model improvement. These findings highlight the importance of advanced red-teaming strategies in ensuring the reliability of VLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- TreeTeaming: Autonomous Red-Teaming of Vision-Language Models via Hierarchical Strategy ExplorationChunxiao Li, Lijun Li, Jing ShaoCVPR 2026 · 被引用 4 次
- TEAR: Temporal-aware Automated Red-teaming for Text-to-Video ModelsJiaming He, Guanyu Hou, Hongwei Li, Zhicong Huang 等CVPR 2026 · 被引用 3 次
- When Search Goes Wrong: Red-Teaming Web-Augmented Large Language ModelsHaoran Ou, Kangjie Chen, Xingshuo Han, Gelei Deng 等ICML 2026 · 被引用 2 次
- VERA-V: Variational Inference Framework for Jailbreaking Vision-Language ModelsQilin Liao, Anamika Lochab, Ruqi ZhangICML 2026 · 被引用 1 次
- STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity AttackXutao Mao, Liangjie Zhao, Tao Liu, Xiang Zheng 等ICML 2026
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
相关 Paper
- Automated Red Teaming for Text-to-Image Models Through Feedback-Guided Prompt Iteration with Vision-Language ModelsWei Xu, Kangjie Chen, Jiawei Qiu, Yuyang Zhang 等ICCV 2025 · 被引用 3 次
- ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play AttacksZhaorun Chen, Xun Liu, Mintong Kang, Jiawei Zhang 等ICLR 2026 · 被引用 5 次
- Backdooring Vision-Language Models with Out-Of-Distribution DataWeimin Lyu, Jiachen Yao, Saumya Gupta, Lu Pang 等ICLR 2025
- Breaking Cross-modal Alignment in Embodied Intelligence: A Multimodal Adversarial Attack Framework for Vision-Language-Action ModelsZhihui Zhao, Xiaorong Dong, Yaowen Zheng, Xiaohui Chen 等WWW 2026
- Towards Adversarial Attack on Vision-Language Pre-training ModelsJiaming Zhang, Qi Yi, Jitao SangACM MM 2022 · 被引用 111 次
