Global-Local Characteristic Excited Cross-Modal Attacks from Images to Videos
Ruikui Wang, Yuanfang Guo, Yunhong Wang
摘要
The transferability of adversarial examples is the key property in practical black-box scenarios. Currently, numerous methods improve the transferability across different models trained on the same modality of data. The investigation of generating video adversarial examples with imagebased substitute models to attack the target video models, i.e., cross-modal transferability of adversarial examples, is rarely explored. A few works on cross-modal transferability directly apply image attack methods for each frame and no factors especial for video data are considered, which limits the cross-modal transferability of adversarial examples. In this paper, we propose an effective cross-modal attack method which considers both the global and local characteristics of video data. Firstly, from the global perspective, we introduce inter-frame interaction into attack process to induce more diverse and stronger gradients rather than perturb each frame separately. Secondly, from the local perspective, we disrupt the inherently local correlation of frames within a video, which prevents black-box video model from capturing valuable temporal clues. Extensive experiments on the UCF-101 and Kinetics-400 validate the proposed method significantly improves cross-modal transferability and even surpasses stronger baseline using video models as substitute model. Our source codes are available at https://github.com/lwmming/Cross-Modal-Attack.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Transferability of Adversarial Attacks in Video-based MLLMs: A Cross-modal Image-to-Video ApproachLinhao Huang, Xue Jiang, Zhiqiang Wang, Wentao Mo 等AAAI 2026 · 被引用 6 次
- ReToMe-VA: Recursive Token Merging for Video Diffusion-based Unrestricted Adversarial AttackZiyi Gao, Kai Chen, Zhipeng Wei, Tingshu Mou 等ACM MM 2024 · 被引用 3 次
- FeatureFool: Zero-Query Fooling of Video Models via Feature MapDuoxun Tang, Xi Xiao, Guangwu Hu, Kangkang Sun 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper16
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- Nesterov Accelerated Gradient and Scale Invariance for Adversarial AttacksJiadong Lin, Chuanbiao Song, Kun He, Liwei Wang 等ICLR 2020 · 被引用 765 次
- Enhancing Adversarial Example Transferability With an Intermediate Level AttackQian Huang, Isay Katsman, Zeqi Gu, Horace He 等ICCV 2019 · 被引用 293 次
相关 Paper
- Cross-Modal Transferable Adversarial Attacks from Images to VideosZhipeng Wei, Jingjing Chen, Zuxuan Wu, Yu-Gang JiangCVPR 2022 · 被引用 45 次
- Boosting the Transferability of Video Adversarial Examples via Temporal TranslationZhipeng Wei, Jingjing Chen, Zuxuan Wu, Yu-Gang JiangAAAI 2022 · 被引用 48 次
- GCMA: Generative Cross-Modal Transferable Adversarial Attacks from Images to VideosKai Chen, Zhipeng Wei, Jingjing Chen, Zuxuan Wu 等ACM MM 2023 · 被引用 13 次
- Heuristic Black-Box Adversarial Attacks on Video Recognition ModelsZhipeng Wei, Jingjing Chen, Xingxing Wei, Linxi Jiang 等AAAI 2020 · 被引用 84 次
- Boosting Adversarial Transferability using Dynamic CuesMuzammal Naseer, Ahmad Mahmood, Salman Khan, Fahad Shahbaz KhanICLR 2023 · 被引用 2 次
