M3Net: Multi-view Encoding, Matching, and Fusion for Few-shot Fine-grained Action Recognition
Hao Tang, Jun Liu, Shuanglin Yan, Rui Yan, Zechao Li, Jinhui Tang
摘要
Due to the scarcity of manually annotated data required for finegrained video understanding, few-shot fine-grained (FS-FG) action recognition has gained significant attention, with the aim of classifying novel fine-grained action categories with only a few labeled instances. Despite the progress made in FS coarse-grained action recognition, current approaches encounter two challenges when dealing with the fine-grained action categories: the inability to capture subtle action details and the insufficiency of learning from limited data that exhibit high intra-class variance and inter-class similarity. To address these limitations, we propose M 3 Net, a matchingbased framework for FS-FG action recognition, which incorporates multi-view encoding, multi-view matching, and multi-view fusion to facilitate embedding encoding, similarity matching, and decision making across multiple viewpoints. Multi-view encoding captures rich contextual details from the intra-frame, intra-video, and intraepisode perspectives, generating customized higher-order embeddings for fine-grained data. Multi-view matching integrates various matching functions enabling flexible relation modeling within limited samples to handle multi-scale spatio-temporal variations by leveraging the instance-specific, category-specific, and task-specific perspectives. Multi-view fusion consists of matching-predictions fusion and matching-losses fusion over the above views, where the former promotes mutual complementarity and the latter enhances embedding generalizability by employing multi-task collaborative learning. Explainable visualizations and experimental results on three challenging benchmarks demonstrate the superiority of M 3 Net in capturing fine-grained action details and achieving state-of-the-art performance for FS-FG action recognition.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Delving into Multimodal Prompting for Fine-Grained Visual ClassificationXin Jiang, Hao Tang, Junyao Gao, Xiaoyu Du 等AAAI 2024 · 被引用 71 次
- Prototypical Prompting for Text-to-image Person Re-identificationShuanglin Yan, Jun Liu, Neng Dong, Liyan Zhang 等ACM MM 2024 · 被引用 16 次
- SeFAR: Semi-supervised Fine-grained Action Recognition with Temporal Perturbation and Learning StabilizationYongle Huang, Haodong Chen, Zhenbang Xu, Zihan Jia 等AAAI 2025 · 被引用 13 次
- Multi-scale Activation, Selection, and Aggregation: Exploring Diverse Cues for Fine-Grained Bird RecognitionZhicheng Zhang, Hao Tang, Jinhui TangAAAI 2025 · 被引用 6 次
- Unveiling the Fragility of Vision-Language Models: Multi-Modal Adversarial Synergy via Texture-Constrained Perturbations and Cross-Modal OptimizationXiang Fang, Wanlong Fang, Changshuo WangAAAI 2026 · 被引用 3 次
它引用的顶会 Paper17
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Causal Intervention for Weakly-Supervised Semantic SegmentationDong Zhang, Hanwang Zhang, Jinhui Tang, Xian-Sheng Hua 等NeurIPS 2020 · 被引用 563 次
- Few-Shot Image Recognition With Knowledge TransferZhimao Peng, Zechao Li, Junge Zhang, Yan Li 等ICCV 2019 · 被引用 230 次
- Spatio-temporal Relation Modeling for Few-shot Action RecognitionAnirudh Thatipelli, Sanath Narayan, Salman Khan, Rao Muhammad Anwer 等CVPR 2022 · 被引用 144 次
- Hybrid Relation Guided Set Matching for Few-shot Action RecognitionXiang Wang, Shiwei Zhang, Zhiwu Qing, Mingqian Tang 等CVPR 2022 · 被引用 124 次
相关 Paper
- MPL: Match-guided Prototype Learning for Few-shot Action RecognitionFeng Yang, Jie Zhao, Fulin Luo, Anyong Qin 等CVPR 2026
- Few-shot Fine-Grained Action Recognition via Bidirectional Attention and Contrastive Meta-LearningJiahao Wang, Yunhong Wang, Sheng Liu, Annan LiACM MM 2021 · 被引用 15 次
- Dual Attention Networks for Few-Shot Fine-Grained RecognitionShu-Lin Xu, Faen Zhang, Xiu-Shen Wei, Jianhua WangAAAI 2022 · 被引用 43 次
- Depth Guided Adaptive Meta-Fusion Network for Few-shot Video RecognitionYuqian Fu, Li Zhang, Junke Wang, Yanwei Fu 等ACM MM 2020 · 被引用 97 次
- Hierarchical Reasoning Network with Contrastive Learning for Few-Shot Human-Object Interaction RecognitionJiale Yu, Baopeng Zhang, Qirui Li, Haoyang Chen 等ACM MM 2023 · 被引用 3 次
