Selective Dependency Aggregation for Action Classification
Yi Tan, Yanbin Hao, Xiangnan He, Yinwei Wei, Xun Yang
Abstract
Video data are distinct from images for the extra temporal dimension, which results in more content dependencies from various perspectives (i.e., long-range and short-range). It increases the difficulty of learning representation for various video actions. Existing methods mainly focus on the dependency under a specific perspective, which cannot facilitate the categorization of complex video actions. This paper proposes a novel selective dependency aggregation (SDA) module, which adaptively exploits multiple types of video dependencies to refine the features. Specifically, we empirically investigate various long-range and short-range dependencies achieved by the multi-direction multi-scale feature squeeze and the dependency excitation. Query structured attention is then adopted to fuse them selectively, fully considering the diversity of videos' dependency preferences. Moreover, the channel reduction mechanism is involved in SDA for controlling the additional computation cost to be lightweight. Finally, we show that the SDA module can be easily plugged into different backbones to form SDA-Nets and demonstrate its effectiveness, efficiency and robustness by conducting extensive experiments on several video benchmarks for action classification. The code and models will be available at https://github.com/ty-97/SDA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Group Contextualization for Video RecognitionYanbin Hao, Hao Zhang, Chong-Wah Ngo, Xiangnan HeCVPR 2022 · 48 citations
- UTC: A Unified Transformer with Inter-Task Contrastive Learning for Visual DialogCheng Chen, Zhenshan Tan, Qingrong Cheng, Xin Jiang et al.CVPR 2022 · 36 citations
- Unified Multi-modal Unsupervised Representation Learning for Skeleton-based Action UnderstandingShengkai Sun, Daizong Liu, Jianfeng Dong, Xiaoye Qu et al.ACM MM 2023 · 34 citations
- Redundancy-aware Transformer for Video Question AnsweringYicong Li, Xun Yang, An Zhang, Chun Feng et al.ACM MM 2023 · 23 citations
- Unsupervised Video Hashing with Multi-granularity Contextualization and Multi-structure PreservationYanbin Hao, Jingru Duan, Hao Zhang, Bin Zhu et al.ACM MM 2022 · 16 citations
Builds on18
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Video Classification With Channel-Separated Convolutional NetworksDu Tran, Heng Wang, Matt Feiszli, Lorenzo TorresaniICCV 2019 · 647 citations
- STM: SpatioTemporal and Motion Encoding for Action RecognitionBoyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu et al.ICCV 2019 · 442 citations
- TAM: Temporal Adaptive Module for Video RecognitionZhaoyang Liu, Limin Wang, Wayne Wu, Chen Qian et al.ICCV 2021 · 356 citations
Related papers
- Colar: Effective and Efficient Online Action Detection by Consulting ExemplarsLe Yang, Junwei Han, Dingwen ZhangCVPR 2022 · 55 citations
- ACTION-Net: Multipath Excitation for Action RecognitionZhengwei Wang, Qi She, Aljosa SmolicCVPR 2021
- TEA: Temporal Excitation and Aggregation for Action RecognitionYan Li, Bin Ji, Xintian Shi, Jianguo Zhang et al.CVPR 2020
- SSAN: Separable Self-Attention Network for Video Representation LearningXudong Guo, Xun Guo, Yan LuCVPR 2021
- Learning Correlation Structures for Vision TransformersManjin Kim, Paul Hongsuck Seo, Cordelia Schmid, Minsu ChoCVPR 2024
