Video-FocalNets: Spatio-Temporal Focal Modulation for Video Action Recognition
Syed Talal Wasim, Muhammad Uzair Khattak, Muzammal Naseer, Salman Khan, Mubarak Shah, Fahad Shahbaz Khan
Abstract
Recent video recognition models utilize Transformer models for long-range spatio-temporal context modeling. Video transformer designs are based on self-attention that can model global context at a high computational cost. In comparison, convolutional designs for videos offer an efficient alternative but lack long-range dependency modeling. Towards achieving the best of both designs, this work proposes Video-FocalNet, an effective and efficient architecture for video recognition that models both local and global contexts. Video-FocalNet is based on a spatiotemporal focal modulation architecture that reverses the interaction and aggregation steps of self-attention for better efficiency. Further, the aggregation step and the interaction step are both implemented using efficient convolution and element-wise multiplication operations that are computationally less expensive than their self-attention counterparts on video representations. We extensively explore the design space of focal modulation-based spatiotemporal context modeling and demonstrate our parallel spatial and temporal encoding design to be the optimal choice. Video-FocalNets perform favorably well against the state-of-the-art transformer-based models for video recognition on five large-scale datasets (Kinetics-400, Kinetics-600, SS-v2, Diving-48, and ActivityNet-1.3) at a lower computational cost. Our code/models are released at https://github.com/TalalWasim/Video-FocalNets .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d01894bd-21fe-498c-8137-e0b63085ff50Cited by top-tier papers8
- BAH Dataset for Ambivalence/Hesitancy Recognition in Videos for Digital Behavioural ChangeManuela González-González, Soufiane Belharbi, Muhammad Osama Zeeshan, Masoumeh Sharafi et al.ICLR 2026 · 18 citations
- SOAP: Enhancing Spatio-Temporal Relation and Motion Information Capturing for Few-Shot Action RecognitionWenbo Huang, Jinghui Zhang, Xuwei Qian, Zhen Wu et al.ACM MM 2024 · 8 citations
- Extending Video Masked Autoencoders to 128 framesNitesh Bharadwaj Gundavarapu, Luke Friedman, Raghav Goyal, Chaitra Hegde et al.NeurIPS 2024 · 6 citations
- Unlocking Motion from Large Vision Models with a Semantic and Kinematic Duality for Gait RecognitionZhanbo Huang, Dingqiang Ye, Xiaoming Liu, Yu KongCVPR 2026 · 4 citations
- Principles of Visual Tokens for Efficient Video UnderstandingXinyue Hao, Gen Li, Shreyank N. Gowda, Robert B. Fisher et al.ICCV 2025 · 1 citation
Builds on33
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
Related papers
- Space-time Mixing Attention for Video TransformerAdrian Bulat, Juan-Manuel Pérez-Rúa, Swathikiran Sudhakaran, Brais Martínez et al.NeurIPS 2021 · 158 citations
- Shrinking Temporal Attention in Transformers for Video Action RecognitionBonan Li, Pengfei Xiong, Congying Han, Tiande GuoAAAI 2022 · 19 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- Adaptive Focus for Efficient Video RecognitionYulin Wang, Zhaoxi Chen, Haojun Jiang, Shiji Song et al.ICCV 2021 · 117 citations
- EAC-Net: Efficient and Accurate Convolutional Network for Video RecognitionBowei Jin, Zhuo XuAAAI 2020 · 2 citations
