Unsupervised Few-Shot Action Recognition via Action-Appearance Aligned Meta-Adaptation
Jay Patravali, Gaurav Mittal, Ye Yu, Fuxin Li, Mei Chen
Abstract
We present MetaUVFS as the first Unsupervised Meta-learning algorithm for Video Few-Shot action recognition. MetaUVFS leverages over 550K unlabeled videos to train a two-stream 2D and 3D CNN architecture via contrastive learning to capture the appearance-specific spatial and action-specific spatio-temporal video features respectively. MetaUVFS comprises a novel Action-Appearance Aligned Meta-adaptation (A3M) module that learns to focus on the action-oriented video features in relation to the appearance features via explicit few-shot episodic meta-learning over unsupervised hard-mined episodes. Our action-appearance alignment and explicit few-shot learner conditions the unsupervised training to mimic the downstream few-shot task, enabling MetaUVFS to significantly outperform all state-of-the-art unsupervised methods on few-shot benchmarks. Moreover, unlike previous few-shot action recognition methods that are supervised, MetaUVFS needs neither base-class labels nor a supervised pretrained backbone. Thus, we need to train MetaUVFS just once to perform competitively or sometimes even outperform state-of-the-art supervised methods on popular HMDB51, UCF101, and Kinetics100 few-shot datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9cd6796a-906e-4525-89f9-bb7b3f6cf5ddCited by top-tier papers4
- Revisiting the Spatial and Temporal Modeling for Few-Shot Action RecognitionJiazheng Xing, Mengmeng Wang, Yong Liu, Boyu MuAAAI 2023 · 51 citations
- Boosting Few-shot Action Recognition with Graph-guided Hybrid MatchingJiazheng Xing, Mengmeng Wang, Yudi Ruan, Bofan Chen et al.ICCV 2023 · 41 citations
- MetaNeRV: Meta Neural Representations for Videos with Spatial-Temporal GuidanceJialong Guo, Ke Liu, Jiangchao Yao, Zhihua Wang et al.AAAI 2025 · 7 citations
- Protect to Adapt: Orthogonal Subspace Control with Ranked Negative-Prompt Curriculum for Few-Shot Action RecognitionHantao Qi, Yan Yan, Junlong Gao, Hanzi WangCVPR 2026
Builds on13
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few ExamplesEleni Triantafillou, Tyler Zhu, Vincent Dumoulin, Pascal Lamblin et al.ICLR 2020 · 692 citations
- Boosting Few-Shot Visual Learning With Self-SupervisionSpyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez et al.ICCV 2019 · 445 citations
- CrossTransformers: spatially-aware few-shot transferCarl Doersch, Ankush Gupta, Andrew ZissermanNeurIPS 2020 · 420 citations
Related papers
- Self-Supervised Video Representation Learning with Meta-Contrastive NetworkYuanze Lin, Xun Guo, Yan LuICCV 2021 · 46 citations
- Depth Guided Adaptive Meta-Fusion Network for Few-shot Video RecognitionYuqian Fu, Li Zhang, Junke Wang, Yanwei Fu et al.ACM MM 2020 · 97 citations
- Meta-GMVAE: Mixture of Gaussian VAE for Unsupervised Meta-LearningDong Bok Lee, Dongchan Min, Seanie Lee, Sung Ju HwangICLR 2021 · 62 citations
- Manta: Enhancing Mamba for Few-Shot Action Recognition of Long Sub-SequenceWenbo Huang, Jinghui Zhang, Guang Li, Lei Zhang et al.AAAI 2025 · 10 citations
- Hierarchical Meta-prototypes Network for Few-shot Action RecognitionXiaoyu Chen, Yigang Cen, Wanru Xu, Yue Zhang et al.ACM MM 2025 · 1 citation
