Few-shot Fine-Grained Action Recognition via Bidirectional Attention and Contrastive Meta-Learning
Jiahao Wang, Yunhong Wang, Sheng Liu, Annan Li
Abstract
Fine-grained action recognition is attracting increasing attention due to the emerging demand of specific action understanding in real-world applications, whereas the data of rare fine-grained categories is very limited. Therefore, we propose the few-shot fine-grained action recognition problem, aiming to recognize novel fine-grained actions with only few samples given for each class. Although progress has been made in coarse-grained actions, existing few-shot recognition methods encounter two issues handling fine-grained actions: the inability to capture subtle action details and the inadequacy in learning from data with low inter-class variance. To tackle the first issue, a human vision inspired bidirectional attention module (BAM) is proposed. Combining top-down task-driven signals with bottom-up salient stimuli, BAM captures subtle action details by accurately highlighting informative spatio-temporal regions. To address the second issue, we introduce contrastive meta-learning (CML). Compared with the widely adopted ProtoNet-based method, CML generates more discriminative video representations for low inter-class variance data, since it makes full use of potential contrastive pairs in each training episode. Furthermore, to fairly compare different models, we establish specific benchmark protocols on two large-scale fine-grained action recognition datasets. Extensive experiments show that our method consistently achieves state-of-the-art performance across evaluated tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79bca012-68ed-4a2e-9943-09a05cd1990dCited by top-tier papers4
- Learning Cross-Image Object Semantic Relation in Transformer for Few-Shot Fine-Grained Image ClassificationBo Zhang, Jiakang Yuan, Baopu Li, Tao Chen et al.ACM MM 2022 · 42 citations
- Exploring Effective Knowledge Transfer for Few-shot Object DetectionZhiyuan Zhao, Qingjie Liu, Yunhong WangACM MM 2022 · 16 citations
- SeFAR: Semi-supervised Fine-grained Action Recognition with Temporal Perturbation and Learning StabilizationYongle Huang, Haodong Chen, Zhenbang Xu, Zihan Jia et al.AAAI 2025 · 13 citations
- Learning Causal Domain-Invariant Temporal Dynamics for Few-Shot Action RecognitionYuke Li, Guangyi Chen, Ben Abramowitz, Stefano Anzellotti et al.ICML 2024 · 3 citations
Builds on16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Motion Guided Attention for Video Salient Object DetectionHaofeng Li, Guanqi Chen, Guanbin Li, Yizhou YuICCV 2019 · 200 citations
Related papers
- Hierarchical Reasoning Network with Contrastive Learning for Few-Shot Human-Object Interaction RecognitionJiale Yu, Baopeng Zhang, Qirui Li, Haoyang Chen et al.ACM MM 2023 · 3 citations
- Dual Attention Networks for Few-Shot Fine-Grained RecognitionShu-Lin Xu, Faen Zhang, Xiu-Shen Wei, Jianhua WangAAAI 2022 · 43 citations
- M3Net: Multi-view Encoding, Matching, and Fusion for Few-shot Fine-grained Action RecognitionHao Tang, Jun Liu, Shuanglin Yan, Rui Yan et al.ACM MM 2023 · 78 citations
- Saliency-Guided Fine-Grained Temporal Mask Learning for Few-Shot Action RecognitionShuo Zheng, Yuanjie Dang, Peng Chen, Ruohong Huan et al.ACM MM 2024 · 1 citation
- Unsupervised Few-Shot Action Recognition via Action-Appearance Aligned Meta-AdaptationJay Patravali, Gaurav Mittal, Ye Yu, Fuxin Li et al.ICCV 2021 · 24 citations
