Breaking Temporal Consistency: Generating Video Universal Adversarial Perturbations Using Image Models
Hee-Seon Kim, Minji Son, Minbeom Kim, Myung-Joon Kwon, Changick Kim
Abstract
As video analysis using deep learning models becomes more widespread, the vulnerability of such models to adversarial attacks is becoming a pressing concern. In particular, Universal Adversarial Perturbation (UAP) poses a significant threat, as a single perturbation can mislead deep learning models on entire datasets. We propose a novel video UAP using image data and image model. This enables us to take advantage of the rich image data and image model-based studies available for video applications. However, there is a challenge that image models are limited in their ability to analyze the temporal aspects of videos, which is crucial for a successful video attack. To address this challenge, we introduce the Breaking Temporal Consistancy (BTC) method, which is the first attempt to incorporate temporal information into video attacks using image models. We aim to generate adversarial videos that have opposite patterns to the original. Specifically, BTC-UAP minimizes the feature similarity between neighboring frames in videos. Our approach is simple but effective at attacking unseen video models. Additionally, it is applicable to videos of varying lengths and invariant to temporal shifts. Our approach surpasses existing methods in terms of effectiveness on various datasets, including ImageNet, UCF-101, and Kinetics-400.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77fe7311-ebc0-4748-bf93-309a04dffa3dCited by top-tier papers3
- From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task KnowledgeHui Lu, Yi Yu, Song Xia, Yiming Yang et al.AAAI 2026 · 8 citations
- Transferability of Adversarial Attacks in Video-based MLLMs: A Cross-modal Image-to-Video ApproachLinhao Huang, Xue Jiang, Zhiqiang Wang, Wentao Mo et al.AAAI 2026 · 6 citations
- Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality PerspectivesZeliang Zhang, Susan Liang, Daiki Shimada, Chenliang XuICLR 2025
Builds on17
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Nesterov Accelerated Gradient and Scale Invariance for Adversarial AttacksJiadong Lin, Chuanbiao Song, Kun He, Liwei Wang et al.ICLR 2020 · 765 citations
- Stealthy Adversarial Perturbations Against Real-Time Video Classification SystemsShasha Li, Ajaya Neupane, Sujoy Paul, Chengyu Song et al.NDSS 2019 · 132 citations
- Universal Adversarial Perturbation via Prior Driven Uncertainty ApproximationHong Liu, Rongrong Ji, Jie Li, Baochang Zhang et al.ICCV 2019 · 90 citations
Related papers
- Boosting the Transferability of Video Adversarial Examples via Temporal TranslationZhipeng Wei, Jingjing Chen, Zuxuan Wu, Yu-Gang JiangAAAI 2022 · 48 citations
- Over-the-Air Adversarial Flickering Attacks Against Video Recognition NetworksRoi Pony, Itay Naeh, Shie MannorCVPR 2021
- GCMA: Generative Cross-Modal Transferable Adversarial Attacks from Images to VideosKai Chen, Zhipeng Wei, Jingjing Chen, Zuxuan Wu et al.ACM MM 2023 · 13 citations
- Universal 3-Dimensional Perturbations for Black-Box Attacks on Video Recognition SystemsShangyu Xie, Han Wang, Yu Kong, Yuan HongS&P 2022 · 32 citations
- Learning Universal Adversarial Perturbation by Adversarial ExampleMaosen Li, Yanhua Yang, Kun Wei, Xu Yang et al.AAAI 2022 · 44 citations
