Searching for Better Spatio-temporal Alignment in Few-Shot Action Recognition
Yichao Cao, Xiu Su, Qingfei Tang, Shan You, Xiaobo Lu, Chang Xu
Abstract
Spatio-Temporal feature matching and alignment are essential for few-shot action recognition as they determine the coherence and effectiveness of the temporal patterns. Nevertheless, this process could be not reliable, especially when dealing with complex video scenarios. In this paper, we propose to improve the performance of matching and alignment from the end-to-end design of models. Our solution comes at two-folds. First, we encourage to enhance the extracted Spatio-Temporal representations from few-shot videos in the perspective of architectures. With this aim, we propose a specialized transformer search method for videos, thus the spatial and temporal attention can be well-organized and optimized for stronger feature representations. Second, we also design an efficient non-parametric spatio-temporal prototype alignment strategy to better handle the high variability of motion. In particular, a query-specific class prototype will be generated for each query sample and category, which can better match query sequences against all support sequences. By doing so, our method SST enjoys significant superiority over the benchmark UCF101 and HMDB51 datasets. For example, with no pretraining, our method achieves 17.1% Top-1 accuracy improvement than the baseline TRX on UCF101 5-way 1-shot setting but with only 3x fewer FLOPs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext de065a3a-0bef-4408-9e8b-f5f4456c8dbbCited by top-tier papers3
- Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation ModelsYichao Cao, Qingfei Tang, Xiu Su, Song Chen et al.NeurIPS 2023 · 64 citations
- Re-mine, Learn and Reason: Exploring the Cross-modal Semantic Correlations for Language-guided HOI detectionYichao Cao, Qingfei Tang, Feng Yang, Xiu Su et al.ICCV 2023 · 31 citations
- Detecting Any instruction-to-answer interaction relationship: Universal Instruction-to-Answer Navigator for Med-VQAZhongze Wu, Hongyan Xu, Yitian Long, Shan You et al.ICML 2024 · 4 citations
Builds on17
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster InferenceBenjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock et al.ICCV 2021 · 1,009 citations
- CrossTransformers: spatially-aware few-shot transferCarl Doersch, Ankush Gupta, Andrew ZissermanNeurIPS 2020 · 420 citations
- AutoFormer: Searching Transformers for Visual RecognitionMinghao Chen, Houwen Peng, Jianlong Fu, Haibin LingICCV 2021 · 335 citations
Related papers
- Temporal-Relational CrossTransformers for Few-Shot Action RecognitionToby Perrett, Alessandro Masullo, Tilo Burghardt, Majid Mirmehdi et al.CVPR 2021
- Boosting Few-shot Action Recognition with Graph-guided Hybrid MatchingJiazheng Xing, Mengmeng Wang, Yudi Ruan, Bofan Chen et al.ICCV 2023 · 41 citations
- Few-Shot Transformation of Common Actions Into Time and SpacePengwan Yang, Pascal Mettes, Cees G. M. SnoekCVPR 2021
- Task-Adapter: Task-specific Adaptation of Image Models for Few-shot Action RecognitionCongqi Cao, Yueran Zhang, Yating Yu, Qinyi Lv et al.ACM MM 2024 · 11 citations
- On the Importance of Spatial Relations for Few-shot Action RecognitionYilun Zhang, Yuqian Fu, Xingjun Ma, Lizhe Qi et al.ACM MM 2023 · 20 citations
