MPL: Match-guided Prototype Learning for Few-shot Action Recognition
Feng Yang, Jie Zhao, Fulin Luo, Anyong Qin, Tiecheng Song, Yue Zhao, Chenqiang Gao, Junwei Han
Abstract
Current few-shot action recognition methods achieve impressive performance by learning representative prototypes and designing diverse video matching strategies. However, these approaches typically face two critical limitations: i) prototypes learned through implicit sample interactions lack clear semantic correspondence between querysupport pairs, limiting their class representativeness; ii) the independent design of prototype learning and matching mechanisms creates a potential incompatibility between prototype representations and matching strategies. To address these limitations, we propose a Match-guided Prototype Learning (MPL) method comprising two key components: enhanced match (E-Match) and key-frame extraction match (K-Match). E-Match explicitly enhances prototype learning in class-specific embeddings by incorporating the matched semantics of query samples, while K-Match further refines the prototype representation through key-frame matching at the fine-grained frame level. Additionally, we propose a Cross-Shot Attention Aggregator (CSA-Aggregator) that dynamically aggregates adjacent frames across support samples, thereby obtaining a prototype representation that captures intra-class shared action patterns. In this way, the proposed MPL effectively mines coarse-to-fine, match-guided semantic information from query-support pairs to generate discriminative class prototypes, and improve the compatibility of prototype representation with the match mechanism. Extensive evaluations on four public datasets confirm that MPL achieves superior performance over leading few-shot action recognition techniques. The source code is available at https: //github.com/jayzh-research/MPL-FSAR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Spatio-temporal Relation Modeling for Few-shot Action RecognitionAnirudh Thatipelli, Sanath Narayan, Salman Khan, Rao Muhammad Anwer et al.CVPR 2022 · 144 citations
- Hybrid Relation Guided Set Matching for Few-shot Action RecognitionXiang Wang, Shiwei Zhang, Zhiwu Qing, Mingqian Tang et al.CVPR 2022 · 124 citations
- TA2N: Two-Stage Action Alignment Network for Few-Shot Action RecognitionShuyuan Li, Huabin Liu, Rui Qian, Yuxi Li et al.AAAI 2022 · 98 citations
- M3Net: Multi-view Encoding, Matching, and Fusion for Few-shot Fine-grained Action RecognitionHao Tang, Jun Liu, Shuanglin Yan, Rui Yan et al.ACM MM 2023 · 78 citations
Related papers
- Boosting Few-shot Action Recognition with Graph-guided Hybrid MatchingJiazheng Xing, Mengmeng Wang, Yudi Ruan, Bofan Chen et al.ICCV 2023 · 41 citations
- Saliency-Guided Fine-Grained Temporal Mask Learning for Few-Shot Action RecognitionShuo Zheng, Yuanjie Dang, Peng Chen, Ruohong Huan et al.ACM MM 2024 · 1 citation
- Multi-Speed Global Contextual Subspace Matching for Few-Shot Action RecognitionTianwei Yu, Peng Chen, Yuanjie Dang, Ruohong Huan et al.ACM MM 2023 · 10 citations
- MoLo: Motion-Augmented Long-Short Contrastive Learning for Few-Shot Action RecognitionXiang Wang, Shiwei Zhang, Zhiwu Qing, Changxin Gao et al.CVPR 2023
- Few-Shot Video Classification via Representation Fusion and Promotion LearningHaifeng Xia, Kai Li, Martin Renqiang Min, Zhengming DingICCV 2023 · 13 citations
