Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment
Xiaojun Jia, Sensen Gao, Simeng Qin, Tianyu Pang, Chao Du, Yihao Huang, Xinfeng Li, Yiming Li, Bo Li, Yang Liu
摘要
Multimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples. While existing methods typically achieve targeted attacks by aligning global features-such as CLIP's [CLS] token-between adversarial and target samples, they often overlook the rich local information encoded in patch tokens. This leads to suboptimal alignment and limited transferability, particularly for closed-source models. To address this limitation, we propose a targeted transferable adversarial attack method based on feature optimal alignment, called FOA-Attack, to improve adversarial transfer capability. Specifically, at the global level, we introduce a global feature loss based on cosine similarity to align the coarse-grained features of adversarial samples with those of target samples. At the local level, given the rich local representations within Transformers, we leverage clustering techniques to extract compact local patterns to alleviate redundant local features. We then formulate local feature alignment between adversarial and target samples as an optimal transport (OT) problem and propose a local clustering optimal transport loss to refine fine-grained feature alignment. Additionally, we propose a dynamic ensemble model weighting strategy to adaptively balance the influence of multiple models during adversarial example generation, thereby further improving transferability. Extensive experiments across various models demonstrate the superiority of the proposed method, outperforming state-of-the-art methods, especially in transferring to closed-source MLLMs. The code is released at https://github.com/jiaxiaojunQAQ/FOA-Attack.
Describe this image.
Two people are riding an elephant through a forested area on a sunny day.
Describe this image. A man and a child are riding an elephant through a lush, wooded area.
An elephant carries a driver and two passengers through a green, outdoor environment.
People are seen riding on the back of an elephant in a natural setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric ReasoningXiang Fang, Wanlong Fang, Changshuo WangCVPR 2026 · 被引用 17 次
- When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action ModelsHui Lu, Yi Yu, Yiming Yang, Chenyu Yi 等CVPR 2026 · 被引用 12 次
- VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language ModelsHefei Mei, Zirui Wang, Shen You, Minjing Dong 等ICLR 2026 · 被引用 9 次
- V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMsSen Nie, Jie Zhang, Jianxin Yan, Shiguang Shan 等CVPR 2026 · 被引用 9 次
- PlugGuard: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk DetectionXiaodan Li, Mengjie Wu, Yao Zhu, Yunna Lv 等ICML 2026 · 被引用 5 次
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
相关 Paper
- Multi-Paradigm Collaborative Adversarial Attack Against Multi-Modal Large Language ModelsYuanbo Li, Tianyang Xu, Cong Hu, Tao Zhou 等CVPR 2026 · 被引用 3 次
- Attacking Gray-Box Large Vision-Language Models with Adaptive SVD-Structured Adversarial AlignmentDaizong Liu, Xiaowen Cai, Junhao Dong, Zhongliang Guo 等ICML 2026
- GAMA: Generative Adversarial Multi-Object Scene AttacksAbhishek Aich, Calvin-Khang Ta, Akash Gupta, Chengyu Song 等NeurIPS 2022 · 被引用 26 次
- FLIP: Cross-domain Face Anti-spoofing with Language GuidanceKoushik Srivatsan, Muzammal Naseer, Karthik NandakumarICCV 2023 · 被引用 84 次
- SABER: Spatially Consistent 3D Universal Adversarial Objects for BEV DetectorsAixuan Li, Mochu Xiang, Bosen Hou, Zhexiong Wan 等CVPR 2026 · 被引用 1 次
