Video-FocalNets: Spatio-Temporal Focal Modulation for Video Action Recognition
Syed Talal Wasim, Muhammad Uzair Khattak, Muzammal Naseer, Salman Khan, Mubarak Shah, Fahad Shahbaz Khan
摘要
Recent video recognition models utilize Transformer models for long-range spatio-temporal context modeling. Video transformer designs are based on self-attention that can model global context at a high computational cost. In comparison, convolutional designs for videos offer an efficient alternative but lack long-range dependency modeling. Towards achieving the best of both designs, this work proposes Video-FocalNet, an effective and efficient architecture for video recognition that models both local and global contexts. Video-FocalNet is based on a spatiotemporal focal modulation architecture that reverses the interaction and aggregation steps of self-attention for better efficiency. Further, the aggregation step and the interaction step are both implemented using efficient convolution and element-wise multiplication operations that are computationally less expensive than their self-attention counterparts on video representations. We extensively explore the design space of focal modulation-based spatiotemporal context modeling and demonstrate our parallel spatial and temporal encoding design to be the optimal choice. Video-FocalNets perform favorably well against the state-of-the-art transformer-based models for video recognition on five large-scale datasets (Kinetics-400, Kinetics-600, SS-v2, Diving-48, and ActivityNet-1.3) at a lower computational cost. Our code/models are released at https://github.com/TalalWasim/Video-FocalNets .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- BAH Dataset for Ambivalence/Hesitancy Recognition in Videos for Digital Behavioural ChangeManuela González-González, Soufiane Belharbi, Muhammad Osama Zeeshan, Masoumeh Sharafi 等ICLR 2026 · 被引用 18 次
- SOAP: Enhancing Spatio-Temporal Relation and Motion Information Capturing for Few-Shot Action RecognitionWenbo Huang, Jinghui Zhang, Xuwei Qian, Zhen Wu 等ACM MM 2024 · 被引用 8 次
- Extending Video Masked Autoencoders to 128 framesNitesh Bharadwaj Gundavarapu, Luke Friedman, Raghav Goyal, Chaitra Hegde 等NeurIPS 2024 · 被引用 6 次
- Unlocking Motion from Large Vision Models with a Semantic and Kinematic Duality for Gait RecognitionZhanbo Huang, Dingqiang Ye, Xiaoming Liu, Yu KongCVPR 2026 · 被引用 4 次
- Principles of Visual Tokens for Efficient Video UnderstandingXinyue Hao, Gen Li, Shreyank N. Gowda, Robert B. Fisher 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper33
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
相关 Paper
- Space-time Mixing Attention for Video TransformerAdrian Bulat, Juan-Manuel Pérez-Rúa, Swathikiran Sudhakaran, Brais Martínez 等NeurIPS 2021 · 被引用 158 次
- Shrinking Temporal Attention in Transformers for Video Action RecognitionBonan Li, Pengfei Xiong, Congying Han, Tiande GuoAAAI 2022 · 被引用 19 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Adaptive Focus for Efficient Video RecognitionYulin Wang, Zhaoxi Chen, Haojun Jiang, Shiji Song 等ICCV 2021 · 被引用 117 次
- EAC-Net: Efficient and Accurate Convolutional Network for Video RecognitionBowei Jin, Zhuo XuAAAI 2020 · 被引用 2 次
