Revisiting the Spatial and Temporal Modeling for Few-Shot Action Recognition
Jiazheng Xing, Mengmeng Wang, Yong Liu, Boyu Mu
Abstract
Spatial and temporal modeling is one of the most core aspects of few-shot action recognition. Most previous works mainly focus on long-term temporal relation modeling based on high-level spatial representations, without considering the crucial low-level spatial features and short-term temporal relations. Actually, the former feature could bring rich local semantic information, and the latter feature could represent motion characteristics of adjacent frames, respectively. In this paper, we propose SloshNet, a new framework that revisits the spatial and temporal modeling for few-shot action recognition in a finer manner. First, to exploit the low-level spatial features, we design a feature fusion architecture search module to automatically search for the best combination of the low-level and high-level spatial features. Next, inspired by the recent transformer, we introduce a long-term temporal modeling module to model the global temporal relations based on the extracted spatial appearance features. Meanwhile, we design another short-term temporal modeling module to encode the motion characteristics between adjacent frame representations. After that, the final predictions can be obtained by feeding the embedded rich spatial-temporal features to a common frame-level class prototype matcher. We extensively validate the proposed SloshNet on four few-shot action recognition datasets, including Something-Something V2, Kinetics, UCF101, and HMDB51. It achieves favorable results against state-of-the-art methods in all datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0dd3c146-c8ed-463d-8ba4-b1fba217be02Cited by top-tier papers9
- Boosting Few-shot Action Recognition with Graph-guided Hybrid MatchingJiazheng Xing, Mengmeng Wang, Yudi Ruan, Bofan Chen et al.ICCV 2023 · 41 citations
- Parallel Attention Interaction Network for Few-Shot Skeleton-based Action RecognitionXingyu Liu, Sanping Zhou, Le Wang, Gang HuaICCV 2023 · 17 citations
- Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action RecognitionBozheng Li, Mushui Liu, Gaoang Wang, Yunlong YuAAAI 2025 · 14 citations
- Task-Adapter: Task-specific Adaptation of Image Models for Few-shot Action RecognitionCongqi Cao, Yueran Zhang, Yating Yu, Qinyi Lv et al.ACM MM 2024 · 11 citations
- SOAP: Enhancing Spatio-Temporal Relation and Motion Information Capturing for Few-Shot Action RecognitionWenbo Huang, Jinghui Zhang, Xuwei Qian, Zhen Wu et al.ACM MM 2024 · 8 citations
Builds on9
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- CrossTransformers: spatially-aware few-shot transferCarl Doersch, Ankush Gupta, Andrew ZissermanNeurIPS 2020 · 420 citations
- Spatio-temporal Relation Modeling for Few-shot Action RecognitionAnirudh Thatipelli, Sanath Narayan, Salman Khan, Rao Muhammad Anwer et al.CVPR 2022 · 144 citations
- TA2N: Two-Stage Action Alignment Network for Few-Shot Action RecognitionShuyuan Li, Huabin Liu, Rui Qian, Yuxi Li et al.AAAI 2022 · 98 citations
- Semantic-Guided Relation Propagation Network for Few-shot Action RecognitionXiao Wang, Weirong Ye, Zhongang Qi, Xun Zhao et al.ACM MM 2021 · 40 citations
Related papers
- Temporal-Relational CrossTransformers for Few-Shot Action RecognitionToby Perrett, Alessandro Masullo, Tilo Burghardt, Majid Mirmehdi et al.CVPR 2021
- Searching for Better Spatio-temporal Alignment in Few-Shot Action RecognitionYichao Cao, Xiu Su, Qingfei Tang, Shan You et al.NeurIPS 2022 · 13 citations
- Shrinking Temporal Attention in Transformers for Video Action RecognitionBonan Li, Pengfei Xiong, Congying Han, Tiande GuoAAAI 2022 · 19 citations
- Saliency-Guided Fine-Grained Temporal Mask Learning for Few-Shot Action RecognitionShuo Zheng, Yuanjie Dang, Peng Chen, Ruohong Huan et al.ACM MM 2024 · 1 citation
- On the Importance of Spatial Relations for Few-shot Action RecognitionYilun Zhang, Yuqian Fu, Xingjun Ma, Lizhe Qi et al.ACM MM 2023 · 20 citations
