Temporal-attentive Covariance Pooling Networks for Video Recognition
Zilin Gao, Qilong Wang, Bingbing Zhang, Qinghua Hu, Peihua Li
Abstract
For video recognition task, a global representation summarizing the whole contents of the video snippets plays an important role for the final performance. However, existing video architectures usually generate it by using a simple, global average pooling (GAP) method, which has limited ability to capture complex dynamics of videos. For image recognition task, there exist evidences showing that covariance pooling has stronger representation ability than GAP. Unfortunately, such plain covariance pooling used in image recognition is an orderless representative, which cannot model spatio-temporal structure inherent in videos. Therefore, this paper proposes a Temporal-attentive Covariance Pooling (TCP), inserted at the end of deep architectures, to produce powerful video representations. Specifically, our TCP first develops a temporal attention module to adaptively calibrate spatio-temporal features for the succeeding covariance pooling, approximatively producing attentive covariance representations. Then, a temporal covariance pooling performs temporal pooling of the attentive covariance representations to characterize both intra-frame correlations and inter-frame cross-correlations of the calibrated features. As such, the proposed TCP can capture complex temporal dynamics. Finally, a fast matrix power normalization is introduced to exploit geometry of covariance representations. Note that our TCP is model-agnostic and can be flexibly integrated into any video architectures, resulting in TCPNet for effective video recognition. The extensive experiments on six benchmarks (e.g., Kinetics, Something-Something V1 and Charades) using various video architectures show our TCPNet is clearly superior to its counterparts, while having strong generalization ability. The source code is publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Revitalizing SVD for Global Covariance Pooling: Halley's Method to Overcome Over-FlatteningJiawei Gu, Ziyue Qiao, Xinming Li, Zechao LiNeurIPS 2025 · 5 citations
- Towards Diverse Perspective Learning with Selection over Multiple Temporal PoolingsJihyeon Seong, Jungmin Kim, Jaesik ChoiAAAI 2024 · 1 citation
- Cov2Pose: Leveraging Spatial Covariance for Direct Manifold-aware 6-DoF Object Pose EstimationNassim Ali Ousalah, Peyman Rostami, Vincent Gaudillière, Emmanuel Koumandakis et al.CVPR 2026 · 1 citation
- Neural Koopman Pooling: Control-Inspired Temporal Dynamics Encoding for Skeleton-Based Action RecognitionXinghan Wang, Xin Xu, Yadong MuCVPR 2023
Builds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- STM: SpatioTemporal and Motion Encoding for Action RecognitionBoyuan Jiang, Mengmeng Wang, Weihao Gan, Wei Wu et al.ICCV 2019 · 442 citations
- TAM: Temporal Adaptive Module for Video RecognitionZhaoyang Liu, Limin Wang, Wayne Wu, Chen Qian et al.ICCV 2021 · 356 citations
Related papers
- Contrast and Order Representations for Video Self-supervised LearningKai Hu, Jie Shao, Yuan Liu, Bhiksha Raj et al.ICCV 2021 · 76 citations
- DSANet: Dynamic Segment Aggregation Network for Video-Level Representation LearningWenhao Wu, Yuxiang Zhao, Yanwu Xu, Xiao Tan et al.ACM MM 2021 · 30 citations
- TDN: Temporal Difference Networks for Efficient Action RecognitionLimin Wang, Zhan Tong, Bin Ji, Gangshan WuCVPR 2021
- VidTr: Video Transformer Without ConvolutionsYanyi Zhang, Xinyu Li, Chunhui Liu, Bing Shuai et al.ICCV 2021 · 224 citations
- SSAN: Separable Self-Attention Network for Video Representation LearningXudong Guo, Xun Guo, Yan LuCVPR 2021
