Group Contextualization for Video Recognition
Yanbin Hao, Hao Zhang, Chong-Wah Ngo, Xiangnan He
Abstract
Learning discriminative representation from the complex spatio-temporal dynamic space is essential for video recognition. On top of those stylized spatio-temporal computational units, further refining the learnt feature with axial contexts is demonstrated to be promising in achieving this goal. However, previous works generally focus on utilizing a single kind of contexts to calibrate entire feature channels and could hardly apply to deal with diverse video activities. The problem can be tackled by using pair-wise spatio-temporal attentions to recompute feature response with cross-axis contexts at the expense of heavy computations. In this paper, we propose an efficient feature refinement method that decomposes the feature channels into several groups and separately refines them with different axial contexts in parallel. We refer this lightweight feature calibration as group contextualization (GC). Specifically, we design a family of efficient element-wise calibrators, i.e., ECal-G/S/T/L, where their axial contexts are information dynamics aggregated from other axes either globally or locally, to contextualize feature channel groups. The GC module can be densely plugged into each residual layer of the off-the-shelf video networks. With little computational overhead, consistent improvement is observed when plugging in GC on different networks. By utilizing calibrators to embed feature with four different kinds of contexts in parallel, the learnt representation is expected to be more resilient to diverse types of activities. On videos with rich temporal variations, empirically GC can boost the performance of 2D-CNN (e.g., TSN and TSM) to a level comparable to the state-of-the-art video networks. Code is available at https://github.com/haoyanbin918/Group- Contextualization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 40aa211a-3013-4c26-98f2-8f899fdf1830Cited by top-tier papers14
- GSRFormer: Grounded Situation Recognition Transformer with Alternate Semantic Attention RefinementZhi-Qi Cheng, Qi Dai, Siyao Li, Teruko Mitamura et al.ACM MM 2022 · 40 citations
- Video-FocalNets: Spatio-Temporal Focal Modulation for Video Action RecognitionSyed Talal Wasim, Muhammad Uzair Khattak, Muzammal Naseer, Salman Khan et al.ICCV 2023 · 39 citations
- DFIL: Deepfake Incremental Learning by Exploiting Domain-invariant Forgery CluesKun Pan, Yifang Yin, Yao Wei, Feng Lin et al.ACM MM 2023 · 35 citations
- Emotion-Prior Awareness Network for Emotional Video CaptioningPeipei Song, Dan Guo, Xun Yang, Shengeng Tang et al.ACM MM 2023 · 29 citations
- Unsupervised Video Hashing with Multi-granularity Contextualization and Multi-structure PreservationYanbin Hao, Jingru Duan, Hao Zhang, Bin Zhu et al.ACM MM 2022 · 16 citations
Builds on24
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun et al.ICCV 2021 · 2,947 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Video Swin TransformerZe Liu, Jia Ning, Yue Cao, Yixuan Wei et al.CVPR 2022 · 1,847 citations
Related papers
- TAda! Temporally-Adaptive Convolutions for Video UnderstandingZiyuan Huang, Shiwei Zhang, Liang Pan, Zhiwu Qing et al.ICLR 2022 · 72 citations
- ACTION-Net: Multipath Excitation for Action RecognitionZhengwei Wang, Qi She, Aljosa SmolicCVPR 2021
- Learning Visual Context for Group Activity RecognitionHangjie Yuan, Dong NiAAAI 2021 · 77 citations
- Channelized Axial Attention - considering Channel Relation within Spatial Attention for Semantic SegmentationYe Huang, Di Kang, Wenjing Jia, Liu Liu et al.AAAI 2022 · 44 citations
- Grouped Spatial-Temporal Aggregation for Efficient Action RecognitionChenxu Luo, Alan L. YuilleICCV 2019 · 170 citations
