Alignment-guided Temporal Attention for Video Action Recognition
Yizhou Zhao, Zhenyang Li, Xun Guo, Yan Lu
摘要
Temporal modeling is crucial for various video learning tasks. Most recent approaches employ either factorized (2D+1D) or joint (3D) spatial-temporal operations to extract temporal contexts from the input frames. While the former is more efficient in computation, the latter often obtains better performance. In this paper, we attribute this to a dilemma between the sufficiency and the efficiency of interactions among various positions in different frames. These interactions affect the extraction of task-relevant information shared among frames. To resolve this issue, we prove that frame-by-frame alignments have the potential to increase the mutual information between frame representations, thereby including more task-relevant information to boost effectiveness. Then we propose Alignment-guided Temporal Attention (ATA) to extend 1-dimensional temporal attention with parameter-free patch-level alignments between neighboring frames. It can act as a general plug-in for image backbones to conduct the action recognition task without any model-specific design. Extensive experiments on multiple benchmarks demonstrate the superiority and generality of our module.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Implicit Temporal Modeling with Learnable Alignment for Video RecognitionShuyuan Tu, Qi Dai, Zuxuan Wu, Zhi-Qi Cheng 等ICCV 2023 · 被引用 63 次
- Jointly Modeling Spatio-Temporal Features of Tactile Signals for Action ClassificationJimmy Lin, Junkai Li, Jiasi Gao, Weizhi Ma 等AAAI 2024 · 被引用 3 次
- Video Language Model Pretraining with Spatio-temporal MaskingYue Wu, Zhaobo Qi, Junshu Sun, Yaowei Wang 等CVPR 2025
- BOE-ViT: Boosting Orientation Estimation with Equivariance in Self-Supervised 3D Subtomogram AlignmentRunmin Jiang, Jackson Daggett, Shriya Pingulkar, Yizhou Zhao 等CVPR 2025
- QDFormer: Towards Robust Audiovisual Segmentation in Complex Environments with Quantization-based Semantic DecompositionXiang Li, Jinglu Wang, Xiaohao Xu, Xiulian Peng 等CVPR 2024
它引用的顶会 Paper17
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- ViViT: A Video Vision TransformerAnurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun 等ICCV 2021 · 被引用 2,947 次
相关 Paper
- SSAN: Separable Self-Attention Network for Video Representation LearningXudong Guo, Xun Guo, Yan LuCVPR 2021
- Gated Spatio-Temporal Attention-Guided Video DeblurringMaitreya Suin, A. N. RajagopalanCVPR 2021
- CSTA: CNN-based Spatiotemporal Attention for Video SummarizationJaewon Son, Jaehun Park, Kwangsu KimCVPR 2024 · 被引用 17 次
- TEINet: Towards an Efficient Architecture for Video RecognitionZhaoyang Liu, Donghao Luo, Yabiao Wang, Limin Wang 等AAAI 2020 · 被引用 267 次
- LGDN: Language-Guided Denoising Network for Video-Language ModelingHaoyu Lu, Mingyu Ding, Nanyi Fei, Yuqi Huo 等NeurIPS 2022 · 被引用 20 次
