SSAN: Separable Self-Attention Network for Video Representation Learning
Xudong Guo, Xun Guo, Yan Lu
摘要
Self-attention has been successfully applied to video representation learning due to the effectiveness of modeling long range dependencies. Existing approaches build the dependencies merely by computing the pairwise correlations along spatial and temporal dimensions simultaneously. However, spatial correlations and temporal correlations represent different contextual information of scenes and temporal reasoning. Intuitively, learning spatial contextual information first will benefit temporal modeling. In this paper, we propose a separable self-attention (SSA) module, which models spatial and temporal correlations sequentially, so that spatial contexts can be efficiently used in temporal modeling. By adding SSA module into 2D CNN, we build a SSA network (SSAN) for video representation learning. On the task of video action recognition, our approach outperforms state-of-the-art methods on Something-Something and Kinetics-400 datasets. Our models often outperform counterparts with shallower network and fewer modalities. We further verify the semantic learning ability of our method in visual-language task of video retrieval, which showcases the homogeneity of video representations and text embeddings. On MSR-VTT and Youcook2 datasets, video representations learnt by SSA significantly improve the state-of-the-art performance. * The work was done when the author was with MSRA as an intern.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Group Contextualization for Video RecognitionYanbin Hao, Hao Zhang, Chong-Wah Ngo, Xiangnan HeCVPR 2022 · 被引用 48 次
- FactorizePhys: Matrix Factorization for Multidimensional Attention in Remote Physiological SensingJitesh Joshi, Sos S. Agaian, Youngjun ChoNeurIPS 2024 · 被引用 32 次
- Alignment-guided Temporal Attention for Video Action RecognitionYizhou Zhao, Zhenyang Li, Xun Guo, Yan LuNeurIPS 2022 · 被引用 24 次
- VideoTitans: Scalable Video Prediction with Integrated Short- and Long-term MemoryYoung-Jae Park, Minseok Seo, Hae-Gon JeonNeurIPS 2025 · 被引用 4 次
- Kronecker Mask and Interpretive Prompts are Language-Action Video LearnersJingyi Yang, Zitong Yu, Xiuming Ni, Jia He 等ICLR 2025
它引用的顶会 Paper9
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi 等ICCV 2019 · 被引用 1,437 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
相关 Paper
- Contrast and Order Representations for Video Self-supervised LearningKai Hu, Jie Shao, Yuan Liu, Bhiksha Raj 等ICCV 2021 · 被引用 76 次
- Shrinking Temporal Attention in Transformers for Video Action RecognitionBonan Li, Pengfei Xiong, Congying Han, Tiande GuoAAAI 2022 · 被引用 19 次
- Learning Comprehensive Motion Representation for Action RecognitionMingyu Wu, Boyuan Jiang, Donghao Luo, Junchi Yan 等AAAI 2021 · 被引用 12 次
- Multi-Group Multi-Attention: Towards Discriminative Spatiotemporal RepresentationZhensheng Shi, Liangjie Cao, Cheng Guan, Ju Liang 等ACM MM 2020 · 被引用 1 次
- TDN: Temporal Difference Networks for Efficient Action RecognitionLimin Wang, Zhan Tong, Bin Ji, Gangshan WuCVPR 2021
