Lite-MKD: A Multi-modal Knowledge Distillation Framework for Lightweight Few-shot Action Recognition
Baolong Liu, Tianyi Zheng, Peng Zheng, Daizong Liu, Xiaoye Qu, Junyu Gao, Jianfeng Dong, Xun Wang
Abstract
Existing few-shot action recognition methods have placed primary focus on improving the recognition accuracy while neglecting another important indicator in practical scenarios, i.e., model efficiency. In this paper, we make the first attempt and propose a Lightweight Multi-modal Knowledge Distillation framework (Lite-MKD) for few-shot action recognition. In this framework, the teacher model conducts multi-modal learning to achieve a comprehensive fusion of the optical flow, depth, and appearance features of human movements, thus achieving a more robust representation of actions. The student model is utilized to learn to recognize actions from the single RGB modality at a lower computational cost under the guidance of the teacher. To fully explore and integrate multi-modal information, a hierarchical Multi-modal Fusion Module (MFM) is introduced in the teacher model. Besides, a multi-level Distinguish-to-Mimic (D2M) knowledge distillation component is proposed for the student model. D2M improves the ability of the student model to mimic the action classification probabilities of the teacher model by enhancing the distinguishability of the student model for different video categories in the support set. Extensive experiments on three action recognition datasets Kinetics, HMDB51, and UCF101 demonstrate our framework's effectiveness and stable generalization ability. With a much more lightweight network for inference, we achieve comparable performance to previous state-of-the-art methods. Our source code is available at https://github.com/HuiGuanLab/Lite-MKD
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get eff4803a-2eb6-41c4-8d5c-48d78acc414bCited by top-tier papers5
- Let All Be Whitened: Multi-Teacher Distillation for Efficient Visual RetrievalZhe Ma, Jianfeng Dong, Shouling Ji, Zhenguang Liu et al.AAAI 2024 · 14 citations
- Manta: Enhancing Mamba for Few-Shot Action Recognition of Long Sub-SequenceWenbo Huang, Jinghui Zhang, Guang Li, Lei Zhang et al.AAAI 2025 · 10 citations
- SOAP: Enhancing Spatio-Temporal Relation and Motion Information Capturing for Few-Shot Action RecognitionWenbo Huang, Jinghui Zhang, Xuwei Qian, Zhen Wu et al.ACM MM 2024 · 8 citations
- Beyond Label Semantics:Language-Guided Action Anatomy for Few-Shot Action RecognitionZefeng Qian, Xincheng Yao, Yifei Huang, Chongyang Zhang et al.ICCV 2025 · 4 citations
- Towards Ship License Plate Recognition in the Wild: A Large Benchmark and Strong BaselineBaolong Liu, Ruiqing Yang, Roukai Huang, Wenhao Xu et al.AAAI 2025 · 1 citation
Related papers
- Active Exploration of Multimodal Complementarity for Few-Shot Action RecognitionYuyang Wanyan, Xiaoshan Yang, Chaofan Chen, Changsheng XuCVPR 2023
- Multimodal Fusion via Teacher-Student Network for Indoor Action RecognitionBruce X. B. Yu, Yan Liu, Keith C. C. ChanAAAI 2021 · 75 citations
- Learning an Augmented RGB Representation with Cross-Modal Knowledge Distillation for Action DetectionRui Dai, Srijan Das, François BrémondICCV 2021 · 50 citations
- Decomposed Cross-Modal Distillation for RGB-based Temporal Action DetectionPilhyeon Lee, Taeoh Kim, Minho Shim, Dongyoon Wee et al.CVPR 2023
- Depth Guided Adaptive Meta-Fusion Network for Few-shot Video RecognitionYuqian Fu, Li Zhang, Junke Wang, Yanwei Fu et al.ACM MM 2020 · 97 citations
