Boosting the Transferability of Video Adversarial Examples via Temporal Translation
Zhipeng Wei, Jingjing Chen, Zuxuan Wu, Yu-Gang Jiang
Abstract
Although deep-learning based video recognition models have achieved remarkable success, they are vulnerable to adversarial examples that are generated by adding humanimperceptible perturbations on clean video samples. As indicated in recent studies, adversarial examples are transferable, which makes it feasible for black-box attacks in real-world applications. Nevertheless, most existing adversarial attack methods have poor transferability when attacking other video models and transfer-based attacks on video models are still unexplored. To this end, we propose to boost the transferability of video adversarial examples for black-box attacks on video recognition models. Through extensive analysis, we discover that different video recognition models rely on different discriminative temporal patterns, leading to the poor transferability of video adversarial examples. This motivates us to introduce a temporal translation attack method, which optimizes the adversarial perturbations over a set of temporal translated video clips. By generating adversarial examples over translated videos, the resulting adversarial examples are less sensitive to temporal patterns existed in the whitebox model being attacked and thus can be better transferred. Extensive experiments on the Kinetics-400 dataset and the UCF-101 dataset demonstrate that our method can significantly boost the transferability of video adversarial examples. For transfer-based attack against video recognition models, it achieves a 61.56% average attack success rate on the Kinetics-400 and 48.60% on the UCF-101. Code is available at https://github.com/zhipeng-wei/TT .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fd68db81-90b4-4d56-afee-ecf904e12514Cited by top-tier papers10
- Cross-Modal Transferable Adversarial Attacks from Images to VideosZhipeng Wei, Jingjing Chen, Zuxuan Wu, Yu-Gang JiangCVPR 2022 · 45 citations
- Efficient Decision-based Black-box Patch Attacks on Video RecognitionKaixun Jiang, Zhaoyu Chen, Hao Huang, Jiafeng Wang et al.ICCV 2023 · 30 citations
- Breaking Temporal Consistency: Generating Video Universal Adversarial Perturbations Using Image ModelsHee-Seon Kim, Minji Son, Minbeom Kim, Myung-Joon Kwon et al.ICCV 2023 · 13 citations
- Long-term Leap Attention, Short-term Periodic Shift for Video ClassificationHao Zhang, Lechao Cheng, Yanbin Hao, Chong-Wah NgoACM MM 2022 · 12 citations
- From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task KnowledgeHui Lu, Yi Yu, Song Xia, Yiming Yang et al.AAAI 2026 · 8 citations
Builds on10
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Nesterov Accelerated Gradient and Scale Invariance for Adversarial AttacksJiadong Lin, Chuanbiao Song, Kun He, Liwei Wang et al.ICLR 2020 · 765 citations
- Skip Connections Matter: On the Transferability of Adversarial Examples Generated with ResNetsDongxian Wu, Yisen Wang, Shu-Tao Xia, James Bailey et al.ICLR 2020 · 357 citations
- Heuristic Black-Box Adversarial Attacks on Video Recognition ModelsZhipeng Wei, Jingjing Chen, Xingxing Wei, Linxi Jiang et al.AAAI 2020 · 84 citations
- Zero-Shot Ingredient Recognition by Multi-Relational Graph Convolutional NetworkJingjing Chen, Liangming Pan, Zhipeng Wei, Xiang Wang et al.AAAI 2020 · 59 citations
Related papers
- Global-Local Characteristic Excited Cross-Modal Attacks from Images to VideosRuikui Wang, Yuanfang Guo, Yunhong WangAAAI 2023 · 15 citations
- Universal 3-Dimensional Perturbations for Black-Box Attacks on Video Recognition SystemsShangyu Xie, Han Wang, Yu Kong, Yuan HongS&P 2022 · 32 citations
- Boosting Adversarial Transferability using Dynamic CuesMuzammal Naseer, Ahmad Mahmood, Salman Khan, Fahad Shahbaz KhanICLR 2023 · 2 citations
- Transferability of Adversarial Attacks in Video-based MLLMs: A Cross-modal Image-to-Video ApproachLinhao Huang, Xue Jiang, Zhiqiang Wang, Wentao Mo et al.AAAI 2026 · 6 citations
- Efficient Sparse Attacks on Videos using Reinforcement LearningHuanqian Yan, Xingxing WeiACM MM 2021 · 19 citations
