M3Net: Multi-view Encoding, Matching, and Fusion for Few-shot Fine-grained Action Recognition
Hao Tang, Jun Liu, Shuanglin Yan, Rui Yan, Zechao Li, Jinhui Tang
Abstract
Due to the scarcity of manually annotated data required for finegrained video understanding, few-shot fine-grained (FS-FG) action recognition has gained significant attention, with the aim of classifying novel fine-grained action categories with only a few labeled instances. Despite the progress made in FS coarse-grained action recognition, current approaches encounter two challenges when dealing with the fine-grained action categories: the inability to capture subtle action details and the insufficiency of learning from limited data that exhibit high intra-class variance and inter-class similarity. To address these limitations, we propose M 3 Net, a matchingbased framework for FS-FG action recognition, which incorporates multi-view encoding, multi-view matching, and multi-view fusion to facilitate embedding encoding, similarity matching, and decision making across multiple viewpoints. Multi-view encoding captures rich contextual details from the intra-frame, intra-video, and intraepisode perspectives, generating customized higher-order embeddings for fine-grained data. Multi-view matching integrates various matching functions enabling flexible relation modeling within limited samples to handle multi-scale spatio-temporal variations by leveraging the instance-specific, category-specific, and task-specific perspectives. Multi-view fusion consists of matching-predictions fusion and matching-losses fusion over the above views, where the former promotes mutual complementarity and the latter enhances embedding generalizability by employing multi-task collaborative learning. Explainable visualizations and experimental results on three challenging benchmarks demonstrate the superiority of M 3 Net in capturing fine-grained action details and achieving state-of-the-art performance for FS-FG action recognition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 32eef96c-542d-4c9a-a7ec-ff5c85494dd4Cited by top-tier papers13
- Delving into Multimodal Prompting for Fine-Grained Visual ClassificationXin Jiang, Hao Tang, Junyao Gao, Xiaoyu Du et al.AAAI 2024 · 71 citations
- Prototypical Prompting for Text-to-image Person Re-identificationShuanglin Yan, Jun Liu, Neng Dong, Liyan Zhang et al.ACM MM 2024 · 16 citations
- SeFAR: Semi-supervised Fine-grained Action Recognition with Temporal Perturbation and Learning StabilizationYongle Huang, Haodong Chen, Zhenbang Xu, Zihan Jia et al.AAAI 2025 · 13 citations
- Multi-scale Activation, Selection, and Aggregation: Exploring Diverse Cues for Fine-Grained Bird RecognitionZhicheng Zhang, Hao Tang, Jinhui TangAAAI 2025 · 6 citations
- Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal OptimizationXiang Fang, Wanlong Fang, Changshuo WangAAAI 2026 · 3 citations
Builds on17
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Causal Intervention for Weakly-Supervised Semantic SegmentationDong Zhang, Hanwang Zhang, Jinhui Tang, Xian-Sheng Hua et al.NeurIPS 2020 · 563 citations
- Few-Shot Image Recognition With Knowledge TransferZhimao Peng, Zechao Li, Junge Zhang, Yan Li et al.ICCV 2019 · 230 citations
- Spatio-temporal Relation Modeling for Few-shot Action RecognitionAnirudh Thatipelli, Sanath Narayan, Salman Khan, Rao Muhammad Anwer et al.CVPR 2022 · 144 citations
- Hybrid Relation Guided Set Matching for Few-shot Action RecognitionXiang Wang, Shiwei Zhang, Zhiwu Qing, Mingqian Tang et al.CVPR 2022 · 124 citations
Related papers
- MPL: Match-guided Prototype Learning for Few-shot Action RecognitionFeng Yang, Jie Zhao, Fulin Luo, Anyong Qin et al.CVPR 2026
- Few-shot Fine-Grained Action Recognition via Bidirectional Attention and Contrastive Meta-LearningJiahao Wang, Yunhong Wang, Sheng Liu, Annan LiACM MM 2021 · 15 citations
- Dual Attention Networks for Few-Shot Fine-Grained RecognitionShu-Lin Xu, Faen Zhang, Xiu-Shen Wei, Jianhua WangAAAI 2022 · 43 citations
- Depth Guided Adaptive Meta-Fusion Network for Few-shot Video RecognitionYuqian Fu, Li Zhang, Junke Wang, Yanwei Fu et al.ACM MM 2020 · 97 citations
- Hierarchical Reasoning Network with Contrastive Learning for Few-Shot Human-Object Interaction RecognitionJiale Yu, Baopeng Zhang, Qirui Li, Haoyang Chen et al.ACM MM 2023 · 3 citations
