DART: Distribution-Aware Adaptive Relational Transfer for Adversarial Attacks against Closed-Source MLLMs
Kaidi Hu, Guancheng Wan, Xiao Luo, Ruigang Yang
摘要
This paper studies the critical problem of targeted adversarial attacks against closed-source MLLMs, which aim to generate highly transferable adversarial samples with open-source MLLMs. Previous approaches typically focus on maximizing the similarity of latent representations between adversarial samples and target samples. However, these approaches could overfit specific target samples with severely limited generalization ability to closed-source MLLMs. Towards this end, we propose a novel approach named Distribution-aware Adaptive Relational Transfer (DART) for adversarial attacks against closed-source MLLMs. The core of our DART is to adopt a statistical lens to characterize the intrinsic semantics of images for more generalized and robust alignment. In particular, each augmented image is considered an example from the intrinsic distribution of the original image. Then, we utilize non-parametric Energy Distance to measure the distribution divergence, which is naturally adopted for the semantic alignment in the hidden space. To further enhance transferability to specific target models, we learn a graph neural network (GNN) to explore the complex relations between source and target MLLMs on transferability and adaptively select surrogate models to maximize transferability across diverse targets. Extensive experiments on benchmark datasets validate the superior robustness and effectiveness of the proposed DART in comparison to various competing baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye 等ICLR 2026 · 被引用 670 次
- On Evaluating Adversarial Robustness of Large Vision-Language ModelsYunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang 等NeurIPS 2023 · 被引用 404 次
- Improving the Transferability of Targeted Adversarial Examples through Object-Based Diverse InputJunyoung Byun, Seungju Cho, Myung-Joon Kwon, Hee-Seon Kim 等CVPR 2022 · 被引用 67 次
- Adversarial Attacks against Closed-Source MLLMs via Feature Optimal AlignmentXiaojun Jia, Sensen Gao, Simeng Qin, Tianyu Pang 等NeurIPS 2025 · 被引用 49 次
- Improving Context Understanding in Multimodal Large Language Models via Multimodal Composition LearningWei Li, Hehe Fan, Yongkang Wong, Yi Yang 等ICML 2024 · 被引用 49 次
相关 Paper
- Fit the Distribution: Cross-Image/Prompt Adversarial Attacks on Multimodal Large Language ModelsHai Yan, Haijian Ma, Xiaowen Cai, Daizong Liu 等NeurIPS 2025 · 被引用 21 次
- Rethinking Adversarial Transferability from a Data Distribution PerspectiveYao Zhu, Jiacheng Sun, Zhenguo LiICLR 2022 · 被引用 49 次
- A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1Zhaoyi Li, Xiaohan Zhao, Dong-Dong Wu, Jiacheng Cui 等NeurIPS 2025 · 被引用 43 次
- On Evaluating the Robustness of Large Vision-Language Models via Untargeted Modality Alignment Breaking Adversarial AttackZhichao Li, Hongshan Yang, Zhibo Wang, Huiyu Xu 等USENIX Security 2026
- Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding SpaceLeo Schwinn, David Dobre, Sophie Xhonneux, Gauthier Gidel 等NeurIPS 2024 · 被引用 113 次
