Finding Achilles' Heel: Adversarial Attack on Multi-modal Action Recognition
Deepak Kumar, Chetan Kumar, Chun-Wei Seah, Siyu Xia, Ming Shao
Abstract
Neural network-based models are notoriously known for their adversarial vulnerability. Recent adversarial machine learning mainly focused on images, where a small perturbation can be simply added to fool the learning model. Very recently, this practice has been explored in human action video attacks by adding perturbation to key frames. Unfortunately, frame selection is usually computationally expensive in run-time, and adding noises to all frames is unrealistic, either. In this paper, we present a novel yet efficient approach to address this issue. Multi-modal video data such as RGB, depth and skeleton data have been widely used for human action modeling, and they have been demonstrated with superior performance than a single modality. Interestingly, we observed that the skeleton data is more "vulnerable" under adversarial attack, and we propose to leverage this "Achilles' Heel" to attack multi-modal video data. In particular, first, an adversarial learning paradigm is designed to perturb skeleton data for a specific action under a black box setting, which highlights how body joints and key segments in videos are subject to attack. Second, we propose a graph attention model to explore the semantics between segments from different modalities and within a modality. Third, the attack will be launched in run-time on all modalities through the learned semantics. The proposed method has been extensively evaluated on multi-modal visual action datasets, including PKU-MMD and NTU-RGB+D to validate its effectiveness.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 899f844a-f724-4a79-8f6c-e5f3d72d4caeCited by top-tier papers4
- Self-supervising Action Recognition by Statistical Moment and Subspace DescriptorsLei Wang, Piotr KoniuszACM MM 2021 · 50 citations
- Quantifying and Enhancing Multi-modal Robustness with Modality PreferenceZequn Yang, Yake Wei, Ce Liang, Di HuICLR 2024 · 27 citations
- Defending Black-Box Skeleton-Based Human Activity ClassifiersHe Wang, Yunfeng Diao, Zichang Tan, Guodong GuoAAAI 2023 · 13 citations
- StyleFool: Fooling Video Classification Systems via Style TransferYuxin Cao, Xi Xiao, Ruoxi Sun, Derui Wang et al.S&P 2023
Related papers
- Adversarial Bone Length Attack on Action RecognitionNariki Tanaka, Hiroshi Kera, Kazuhiko KawamotoAAAI 2022 · 18 citations
- Hard No-Box Adversarial Attack on Skeleton-Based Human Action Recognition with Skeleton-Motion-Informed GradientZhengzhi Lu, He Wang, Ziyi Chang, Guoan Yang et al.ICCV 2023 · 17 citations
- Understanding the Robustness of Skeleton-Based Action Recognition Under Adversarial AttackHe Wang, Feixiang He, Zhexi Peng, Tianjia Shao et al.CVPR 2021
- Skeletal Spatial-Temporal Semantics Guided Homogeneous-Heterogeneous Multimodal Network for Action RecognitionChenwei Zhang, Yuxuan Hu, Min Yang, Chengming Li et al.ACM MM 2023 · 4 citations
- Just One Moment: Structural Vulnerability of Deep Action Recognition against One Frame AttackJaehui Hwang, Jun-Hyuk Kim, Jun-Ho Choi, Jong-Seok LeeICCV 2021 · 24 citations
