Searching for Better Spatio-temporal Alignment in Few-Shot Action Recognition
Yichao Cao, Xiu Su, Qingfei Tang, Shan You, Xiaobo Lu, Chang Xu
摘要
Spatio-Temporal feature matching and alignment are essential for few-shot action recognition as they determine the coherence and effectiveness of the temporal patterns. Nevertheless, this process could be not reliable, especially when dealing with complex video scenarios. In this paper, we propose to improve the performance of matching and alignment from the end-to-end design of models. Our solution comes at two-folds. First, we encourage to enhance the extracted Spatio-Temporal representations from few-shot videos in the perspective of architectures. With this aim, we propose a specialized transformer search method for videos, thus the spatial and temporal attention can be well-organized and optimized for stronger feature representations. Second, we also design an efficient non-parametric spatio-temporal prototype alignment strategy to better handle the high variability of motion. In particular, a query-specific class prototype will be generated for each query sample and category, which can better match query sequences against all support sequences. By doing so, our method SST enjoys significant superiority over the benchmark UCF101 and HMDB51 datasets. For example, with no pretraining, our method achieves 17.1% Top-1 accuracy improvement than the baseline TRX on UCF101 5-way 1-shot setting but with only 3x fewer FLOPs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation ModelsYichao Cao, Qingfei Tang, Xiu Su, Song Chen 等NeurIPS 2023 · 被引用 64 次
- Re-mine, Learn and Reason: Exploring the Cross-modal Semantic Correlations for Language-guided HOI detectionYichao Cao, Qingfei Tang, Feng Yang, Xiu Su 等ICCV 2023 · 被引用 31 次
- Detecting Any instruction-to-answer interaction relationship: Universal Instruction-to-Answer Navigator for Med-VQAZhongze Wu, Hongyan Xu, Yitian Long, Shan You 等ICML 2024 · 被引用 4 次
它引用的顶会 Paper17
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- LeViT: a Vision Transformer in ConvNet's Clothing for Faster InferenceBenjamin Graham, Alaaeldin El-Nouby, Hugo Touvron, Pierre Stock 等ICCV 2021 · 被引用 1,009 次
- CrossTransformers: spatially-aware few-shot transferCarl Doersch, Ankush Gupta, Andrew ZissermanNeurIPS 2020 · 被引用 420 次
- AutoFormer: Searching Transformers for Visual RecognitionMinghao Chen, Houwen Peng, Jianlong Fu, Haibin LingICCV 2021 · 被引用 335 次
相关 Paper
- Temporal-Relational CrossTransformers for Few-Shot Action RecognitionToby Perrett, Alessandro Masullo, Tilo Burghardt, Majid Mirmehdi 等CVPR 2021
- Boosting Few-shot Action Recognition with Graph-guided Hybrid MatchingJiazheng Xing, Mengmeng Wang, Yudi Ruan, Bofan Chen 等ICCV 2023 · 被引用 41 次
- Few-Shot Transformation of Common Actions Into Time and SpacePengwan Yang, Pascal Mettes, Cees G. M. SnoekCVPR 2021
- Task-Adapter: Task-specific Adaptation of Image Models for Few-shot Action RecognitionCongqi Cao, Yueran Zhang, Yating Yu, Qinyi Lv 等ACM MM 2024 · 被引用 11 次
- On the Importance of Spatial Relations for Few-shot Action RecognitionYilun Zhang, Yuqian Fu, Xingjun Ma, Lizhe Qi 等ACM MM 2023 · 被引用 20 次
