MPL: Match-guided Prototype Learning for Few-shot Action Recognition
Feng Yang, Jie Zhao, Fulin Luo, Anyong Qin, Tiecheng Song, Yue Zhao, Chenqiang Gao, Junwei Han
摘要
Current few-shot action recognition methods achieve impressive performance by learning representative prototypes and designing diverse video matching strategies. However, these approaches typically face two critical limitations: i) prototypes learned through implicit sample interactions lack clear semantic correspondence between querysupport pairs, limiting their class representativeness; ii) the independent design of prototype learning and matching mechanisms creates a potential incompatibility between prototype representations and matching strategies. To address these limitations, we propose a Match-guided Prototype Learning (MPL) method comprising two key components: enhanced match (E-Match) and key-frame extraction match (K-Match). E-Match explicitly enhances prototype learning in class-specific embeddings by incorporating the matched semantics of query samples, while K-Match further refines the prototype representation through key-frame matching at the fine-grained frame level. Additionally, we propose a Cross-Shot Attention Aggregator (CSA-Aggregator) that dynamically aggregates adjacent frames across support samples, thereby obtaining a prototype representation that captures intra-class shared action patterns. In this way, the proposed MPL effectively mines coarse-to-fine, match-guided semantic information from query-support pairs to generate discriminative class prototypes, and improve the compatibility of prototype representation with the match mechanism. Extensive evaluations on four public datasets confirm that MPL achieves superior performance over leading few-shot action recognition techniques. The source code is available at https: //github.com/jayzh-research/MPL-FSAR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Spatio-temporal Relation Modeling for Few-shot Action RecognitionAnirudh Thatipelli, Sanath Narayan, Salman Khan, Rao Muhammad Anwer 等CVPR 2022 · 被引用 144 次
- Hybrid Relation Guided Set Matching for Few-shot Action RecognitionXiang Wang, Shiwei Zhang, Zhiwu Qing, Mingqian Tang 等CVPR 2022 · 被引用 124 次
- TA2N: Two-Stage Action Alignment Network for Few-Shot Action RecognitionShuyuan Li, Huabin Liu, Rui Qian, Yuxi Li 等AAAI 2022 · 被引用 98 次
- M3Net: Multi-view Encoding, Matching, and Fusion for Few-shot Fine-grained Action RecognitionHao Tang, Jun Liu, Shuanglin Yan, Rui Yan 等ACM MM 2023 · 被引用 78 次
相关 Paper
- Boosting Few-shot Action Recognition with Graph-guided Hybrid MatchingJiazheng Xing, Mengmeng Wang, Yudi Ruan, Bofan Chen 等ICCV 2023 · 被引用 41 次
- Saliency-Guided Fine-Grained Temporal Mask Learning for Few-Shot Action RecognitionShuo Zheng, Yuanjie Dang, Peng Chen, Ruohong Huan 等ACM MM 2024 · 被引用 1 次
- Multi-Speed Global Contextual Subspace Matching for Few-Shot Action RecognitionTianwei Yu, Peng Chen, Yuanjie Dang, Ruohong Huan 等ACM MM 2023 · 被引用 10 次
- MoLo: Motion-Augmented Long-Short Contrastive Learning for Few-Shot Action RecognitionXiang Wang, Shiwei Zhang, Zhiwu Qing, Changxin Gao 等CVPR 2023
- Few-Shot Video Classification via Representation Fusion and Promotion LearningHaifeng Xia, Kai Li, Martin Renqiang Min, Zhengming DingICCV 2023 · 被引用 13 次
