Towards Decision-based Sparse Attacks on Video Recognition
Kaixun Jiang, Zhaoyu Chen, Xinyu Zhou, Jingyu Zhang, Lingyi Hong, Jiafeng Wang, Bo Li, Yan Wang, Wenqiang Zhang
Abstract
Recent studies indicate that sparse attacks threaten the security of deep learning models, which modify only a small set of pixels in the input based on the l0 norm constraint. While existing research has primarily focused on sparse attacks against image models, there is a notable gap in evaluating the robustness of video recognition models. To bridge this gap, we are the first to study sparse video attacks and propose an attack framework named V-DSA in the most challenging decision-based setting, in which threat models only return the predicted hard label. Specifically, V-DSA comprises two modules: a Cross-Modal Generator (CMG) for query-free transfer attacks on each frame and an Optical flow Grouping Evolution algorithm (OGE) for query-efficient spatial-temporal attacks. CMG passes each frame to generate the transfer video as the starting point of the attack based on the feature similarity between image classification and video recognition models. OGE first initializes populations based on transfer video and then leverages optical flow to establish the temporal connection of the perturbed pixels in each frame, which can reduce the parameter space and break the temporal relationship between frames specifically. Finally, OGE complements the above optical flow modeling by grouping evolution which can realize the coarse-to-fine attack to avoid falling into the local optimum. In addition, OGE makes the perturbation with temporal coherence while balancing the number of perturbed pixels per frame, further increasing the imperceptibility of the attack. Extensive experiments demonstrate that V-DSA achieves state-of-the-art performance in terms of both threat effectiveness and imperceptibility. We hope V-DSA can provide valuable insights into the security of video recognition systems.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ae6fc1a1-db83-4642-812d-e32cc484f25cCited by top-tier papers1
Ask how each one uses itRelated papers
- Transferable Structural Sparse Adversarial Attack Via Exact Group Sparsity TrainingDi Ming, Peng Ren, Yunlong Wang, Xin FengCVPR 2024 · 7 citations
- Efficient Decision-based Black-box Patch Attacks on Video RecognitionKaixun Jiang, Zhaoyu Chen, Hao Huang, Jiafeng Wang et al.ICCV 2023 · 30 citations
- Query Efficient Decision Based Sparse Attacks Against Black-Box Deep Learning ModelsViet Quoc Vo, Ehsan Abbasnejad, Damith RanasingheICLR 2022 · 15 citations
- GCMA: Generative Cross-Modal Transferable Adversarial Attacks from Images to VideosKai Chen, Zhipeng Wei, Jingjing Chen, Zuxuan Wu et al.ACM MM 2023 · 13 citations
- Frequency Domain Distributed Perturbations: Towards Query-Efficient Black-Box Adversarial Video AttackTeng Jin, Ziwen He, Zhangjie Fu, Songping Wang et al.ACM MM 2025
