GCMA: Generative Cross-Modal Transferable Adversarial Attacks from Images to Videos
Kai Chen, Zhipeng Wei, Jingjing Chen, Zuxuan Wu, Yu-Gang Jiang
Abstract
Existing cross-domain transferable attacks mostly focus on exploring the adversarial transferability across homomodal domains, while the adversarial transferability across heteromodal domains, e.g., image domains to video domains, has received less attention. This paper investigates cross-modal transferable attacks from image domains to video domains with the generator-oriented approach, i.e., crafting adversarial perturbations for each frame of video clips with the perturbation generator trained in the ImageNet domain to attack target video models. To this end, we propose an effective Generative Cross-Modal Attacks (GCMA) framework to enhance adversarial transferability from image domains to video domains. To narrow the domain gap between image and video data, we first propose a random motion module that warps images with synthetic random optical flows. We then integrate the random motion module into the feature disruption loss to incorporate additional temporal cues in the training phase. Specifically, feature disruption loss minimizes the cosine similarity between intermediate features of warped benign and adversarial images. Furthermore, motivated by the positive correlation between transferability and temporal consistency of adversarial video clips, we also introduce a temporal consistency loss that maximizes the cosine similarity between intermediate features of warped adversarial images and adversarial counterparts of warped benign images. Finally, GCMA trains the perturbation generator by simultaneously optimizing feature disruption loss and temporal consistency loss. Extensive experiments demonstrate the effectiveness of our proposed method, achieving state-of-the-art performance on Kinetics-400 and UCF-101. Our code is available at https://github.com/kay-ck/GCMA.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers2
- CURE: Concept Unlearning via Orthogonal Representation Editing in Diffusion ModelsShristi Das Biswas, Arani Roy, Kaushik RoyNeurIPS 2025 · 30 citations
- AIM: Additional Image Guided Generation of Transferable Adversarial AttacksTeng Li, Xingjun Ma, Yu-Gang JiangAAAI 2025 · 7 citations
Related papers
- Global-Local Characteristic Excited Cross-Modal Attacks from Images to VideosRuikui Wang, Yuanfang Guo, Yunhong WangAAAI 2023 · 15 citations
- Cross-Modal Transferable Adversarial Attacks from Images to VideosZhipeng Wei, Jingjing Chen, Zuxuan Wu, Yu-Gang JiangCVPR 2022 · 45 citations
- Breaking Temporal Consistency: Generating Video Universal Adversarial Perturbations Using Image ModelsHee-Seon Kim, Minji Son, Minbeom Kim, Myung-Joon Kwon et al.ICCV 2023 · 13 citations
- Boosting the Transferability of Video Adversarial Examples via Temporal TranslationZhipeng Wei, Jingjing Chen, Zuxuan Wu, Yu-Gang JiangAAAI 2022 · 48 citations
- From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task KnowledgeHui Lu, Yi Yu, Song Xia, Yiming Yang et al.AAAI 2026 · 8 citations
