Few-Shot Video Classification via Representation Fusion and Promotion Learning
Haifeng Xia, Kai Li, Martin Renqiang Min, Zhengming Ding
Abstract
Recent few-shot video classification (FSVC) works achieve promising performance by capturing similarity across support and query samples with different temporal alignment strategies or learning discriminative features via Transformer block within each episode. However, they ignore two important issues: a) It is difficult to capture rich intrinsic action semantics from a limited number of support instances within each task. b) Redundant or irrelevant frames in videos easily weaken the positive influence of discriminative frames. To address these two issues, this paper proposes a novel Representation Fusion and Promotion Learning (RFPL) mechanism with two sub-modules: meta-action learning (MAL) and reinforced image representation (RIR). Concretely, during training stage, we perform online learning for seeking a task-shared meta-action bank to enrich task-specific action representation by injecting global knowledge. Besides, we exploit reinforcement learning to obtain the importance of each frame and refine the representation. This operation maximizes the contribution of discriminative frames to further capture the similarity of support and query samples from the same category. Our RFPL framework is highly flexible that it can be integrated with many existing FSVC methods. Extensive experiments show that RFPL significantly enhances the performance of existing FSVC models when integrated with them.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Beyond Label Semantics:Language-Guided Action Anatomy for Few-Shot Action RecognitionZefeng Qian, Xincheng Yao, Yifei Huang, Chongyang Zhang et al.ICCV 2025 · 4 citations
- D2 ST-Adapter: Disentangled-and-Deformable Spatio-Temporal Adapter for Few-Shot Action RecognitionWenjie Pei, Qizhong Tan, Guangming Lu, Jiandong Tian et al.ICCV 2025 · 2 citations
- Trokens: Semantic-Aware Relational Trajectory Tokens for Few-Shot Action RecognitionPulkit Kumar, Shuaiyi Huang, Matthew Walmer, Sai Saketh Rambhatla et al.ICCV 2025
- MPL: Match-guided Prototype Learning for Few-shot Action RecognitionFeng Yang, Jie Zhao, Fulin Luo, Anyong Qin et al.CVPR 2026
Builds on13
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Spatio-temporal Relation Modeling for Few-shot Action RecognitionAnirudh Thatipelli, Sanath Narayan, Salman Khan, Rao Muhammad Anwer et al.CVPR 2022 · 144 citations
- Hybrid Relation Guided Set Matching for Few-shot Action RecognitionXiang Wang, Shiwei Zhang, Zhiwu Qing, Mingqian Tang et al.CVPR 2022 · 124 citations
- Generalized and Discriminative Few-Shot Object Detection via SVD-Dictionary EnhancementAming Wu, Suqi Zhao, Cheng Deng, Wei LiuNeurIPS 2021 · 52 citations
Related papers
- Motion-modulated Temporal Fragment Alignment Network For Few-Shot Action RecognitionJiamin Wu, Tianzhu Zhang, Zhe Zhang, Feng Wu et al.CVPR 2022 · 73 citations
- Exploring Stable Meta-Optimization Patterns via Differentiable Reinforcement Learning for Few-Shot ClassificationZheng Han, Xiaobin Zhu, Chun Yang, Hongyang Zhou et al.ACM MM 2024 · 1 citation
- Semantic-Guided Relation Propagation Network for Few-shot Action RecognitionXiao Wang, Weirong Ye, Zhongang Qi, Xun Zhao et al.ACM MM 2021 · 40 citations
- Depth Guided Adaptive Meta-Fusion Network for Few-shot Video RecognitionYuqian Fu, Li Zhang, Junke Wang, Yanwei Fu et al.ACM MM 2020 · 97 citations
- Task-adaptive Spatial-Temporal Video Sampler for Few-shot Action RecognitionHuabin Liu, Weixian Lv, John See, Weiyao LinACM MM 2022 · 11 citations
