Depth Guided Adaptive Meta-Fusion Network for Few-shot Video Recognition
Yuqian Fu, Li Zhang, Junke Wang, Yanwei Fu, Yu-Gang Jiang
Abstract
Humans can easily recognize actions with only a few examples given, while the existing video recognition models still heavily rely on the large-scale labeled data inputs. This observation has motivated an increasing interest in few-shot video action recognition, which aims at learning new actions with only very few labeled samples. In this paper, we propose a depth guided Adaptive Meta-Fusion Network for few-shot video recognition which is termed as AMeFu-Net. Concretely, we tackle the few-shot recognition problem from three aspects: firstly, we alleviate this extremely data-scarce problem by introducing depth information as a carrier of the scene, which will bring extra visual information to our model; secondly, we fuse the representation of original RGB clips with multiple non-strictly corresponding depth clips sampled by our temporal asynchronization augmentation mechanism, which synthesizes new instances at feature-level; thirdly, a novel Depth Guided Adaptive Instance Normalization (DGAdaIN) fusion module is proposed to fuse the two-stream modalities efficiently. Additionally, to better mimic the few-shot recognition process, our model is trained in the meta-learning way. Extensive experiments on several action recognition benchmarks demonstrate the effectiveness of our model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46382af9-0902-45b0-9449-20da9cf41e7eCited by top-tier papers25
- Hybrid Relation Guided Set Matching for Few-shot Action RecognitionXiang Wang, Shiwei Zhang, Zhiwu Qing, Mingqian Tang et al.CVPR 2022 · 124 citations
- M3Net: Multi-view Encoding, Matching, and Fusion for Few-shot Fine-grained Action RecognitionHao Tang, Jun Liu, Shuanglin Yan, Rui Yan et al.ACM MM 2023 · 78 citations
- Motion-modulated Temporal Fragment Alignment Network For Few-Shot Action RecognitionJiamin Wu, Tianzhu Zhang, Zhe Zhang, Feng Wu et al.CVPR 2022 · 73 citations
- Boosting Few-shot Action Recognition with Graph-guided Hybrid MatchingJiazheng Xing, Mengmeng Wang, Yudi Ruan, Bofan Chen et al.ICCV 2023 · 41 citations
- Counterfactual Debiasing Inference for Compositional Action RecognitionPengzhan Sun, Bo Wu, Xunsong Li, Wen Li et al.ACM MM 2021 · 25 citations
Builds on6
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li et al.AAAI 2020 · 4,134 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Instance Credibility Inference for Few-Shot LearningYikai Wang, Chengming Xu, Chen Liu, Li Zhang et al.CVPR 2020
Related papers
- Unsupervised Few-Shot Action Recognition via Action-Appearance Aligned Meta-AdaptationJay Patravali, Gaurav Mittal, Ye Yu, Fuxin Li et al.ICCV 2021 · 24 citations
- Annotation-Efficient Untrimmed Video Action RecognitionYixiong Zou, Shanghang Zhang, Guangyao Chen, Yonghong Tian et al.ACM MM 2021 · 7 citations
- Active Exploration of Multimodal Complementarity for Few-Shot Action RecognitionYuyang Wanyan, Xiaoshan Yang, Chaofan Chen, Changsheng XuCVPR 2023
- Lite-MKD: A Multi-modal Knowledge Distillation Framework for Lightweight Few-shot Action RecognitionBaolong Liu, Tianyi Zheng, Peng Zheng, Daizong Liu et al.ACM MM 2023 · 13 citations
- Semantic-Guided Relation Propagation Network for Few-shot Action RecognitionXiao Wang, Weirong Ye, Zhongang Qi, Xun Zhao et al.ACM MM 2021 · 40 citations
