Depth Guided Adaptive Meta-Fusion Network for Few-shot Video Recognition
Yuqian Fu, Li Zhang, Junke Wang, Yanwei Fu, Yu-Gang Jiang
摘要
Humans can easily recognize actions with only a few examples given, while the existing video recognition models still heavily rely on the large-scale labeled data inputs. This observation has motivated an increasing interest in few-shot video action recognition, which aims at learning new actions with only very few labeled samples. In this paper, we propose a depth guided Adaptive Meta-Fusion Network for few-shot video recognition which is termed as AMeFu-Net. Concretely, we tackle the few-shot recognition problem from three aspects: firstly, we alleviate this extremely data-scarce problem by introducing depth information as a carrier of the scene, which will bring extra visual information to our model; secondly, we fuse the representation of original RGB clips with multiple non-strictly corresponding depth clips sampled by our temporal asynchronization augmentation mechanism, which synthesizes new instances at feature-level; thirdly, a novel Depth Guided Adaptive Instance Normalization (DGAdaIN) fusion module is proposed to fuse the two-stream modalities efficiently. Additionally, to better mimic the few-shot recognition process, our model is trained in the meta-learning way. Extensive experiments on several action recognition benchmarks demonstrate the effectiveness of our model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Hybrid Relation Guided Set Matching for Few-shot Action RecognitionXiang Wang, Shiwei Zhang, Zhiwu Qing, Mingqian Tang 等CVPR 2022 · 被引用 124 次
- M3Net: Multi-view Encoding, Matching, and Fusion for Few-shot Fine-grained Action RecognitionHao Tang, Jun Liu, Shuanglin Yan, Rui Yan 等ACM MM 2023 · 被引用 78 次
- Motion-modulated Temporal Fragment Alignment Network For Few-Shot Action RecognitionJiamin Wu, Tianzhu Zhang, Zhe Zhang, Feng Wu 等CVPR 2022 · 被引用 73 次
- Boosting Few-shot Action Recognition with Graph-guided Hybrid MatchingJiazheng Xing, Mengmeng Wang, Yudi Ruan, Bofan Chen 等ICCV 2023 · 被引用 41 次
- Counterfactual Debiasing Inference for Compositional Action RecognitionPengzhan Sun, Bo Wu, Xunsong Li, Wen Li 等ACM MM 2021 · 被引用 25 次
它引用的顶会 Paper6
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li 等AAAI 2020 · 被引用 4,134 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Instance Credibility Inference for Few-Shot LearningYikai Wang, Chengming Xu, Chen Liu, Li Zhang 等CVPR 2020
相关 Paper
- Unsupervised Few-Shot Action Recognition via Action-Appearance Aligned Meta-AdaptationJay Patravali, Gaurav Mittal, Ye Yu, Fuxin Li 等ICCV 2021 · 被引用 24 次
- Annotation-Efficient Untrimmed Video Action RecognitionYixiong Zou, Shanghang Zhang, Guangyao Chen, Yonghong Tian 等ACM MM 2021 · 被引用 7 次
- Active Exploration of Multimodal Complementarity for Few-Shot Action RecognitionYuyang Wanyan, Xiaoshan Yang, Chaofan Chen, Changsheng XuCVPR 2023
- Lite-MKD: A Multi-modal Knowledge Distillation Framework for Lightweight Few-shot Action RecognitionBaolong Liu, Tianyi Zheng, Peng Zheng, Daizong Liu 等ACM MM 2023 · 被引用 13 次
- Semantic-Guided Relation Propagation Network for Few-shot Action RecognitionXiao Wang, Weirong Ye, Zhongang Qi, Xun Zhao 等ACM MM 2021 · 被引用 40 次
