Multi-Modality Co-Learning for Efficient Skeleton-based Action Recognition
Jinfu Liu, Chen Chen, Mengyuan Liu
Abstract
Skeleton-based action recognition has garnered significant attention due to the utilization of concise and resilient skeletons. Nevertheless, the absence of detailed body information in skeletons restricts performance, while other multimodal methods require substantial inference resources and are inefficient when using multimodal data during both training and inference stages. To address this and fully harness the complementary multimodal features, we propose a novel multi-modality co-learning (MMCL) framework by leveraging the multimodal large language models (LLMs) as auxiliary networks for efficient skeleton-based action recognition, which engages in multi-modality co-learning during the training stage and keeps efficiency by employing only concise skeletons in inference. Our MMCL framework primarily consists of two modules. First, the Feature Alignment Module (FAM) extracts rich RGB features from video frames and aligns them with global skeleton features via contrastive learning. Second, the Feature Refinement Module (FRM) uses RGB images with temporal information and text instruction to generate instructive features based on the powerful generalization of multimodal LLMs. These instructive text features will further refine the classification scores and the refined scores will enhance the model's robustness and generalization in a manner similar to soft labels. Extensive experiments on NTU RGB+D, NTU RGB+D 120 and Northwestern-UCLA benchmarks consistently verify the effectiveness of our MMCL, which outperforms the existing skeleton-based action recognition methods. Meanwhile, experiments on UTD-MHAD and SYSU-Action datasets demonstrate the commendable generalization of our MMCL in zero-shot and domain-adaptive action recognition. Our code is publicly available at: https://github.com/liujf69/MMCL-Action.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action RecognitionYuhang Wen, Mengyuan Liu, Songtao Wu, Beichen DingNeurIPS 2024 · 7 citations
- Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral CloningXiuxiu Qi, Yu Yang, Jiannong Cao, Luyao Bai et al.AAAI 2026 · 2 citations
- Neuron: Learning Context-Aware Evolving Representations for Zero-Shot Skeleton Action RecognitionYang Chen, Jingcai Guo, Song Guo, Dacheng TaoCVPR 2025
- Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-IdentificationRifen Lin, Alex Jinpeng Wang, Jiawei Mo, Min LiAAAI 2026
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionYuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li et al.ICCV 2021 · 871 citations
Related papers
- LLMs are Good Action RecognizersHaoxuan Qu, Yujun Cai, Jun LiuCVPR 2024 · 37 citations
- Generative Action Description Prompts for Skeleton-based Action RecognitionWangmeng Xiang, Chao Li, Yuxuan Zhou, Biao Wang et al.ICCV 2023 · 84 citations
- Unified Multi-modal Unsupervised Representation Learning for Skeleton-based Action UnderstandingShengkai Sun, Daizong Liu, Jianfeng Dong, Xiaoye Qu et al.ACM MM 2023 · 34 citations
- Skeletal Spatial-Temporal Semantics Guided Homogeneous-Heterogeneous Multimodal Network for Action RecognitionChenwei Zhang, Yuxuan Hu, Min Yang, Chengming Li et al.ACM MM 2023 · 4 citations
- Boosting Skeleton-based Zero-Shot Action Recognition with Training-Free Test-Time AdaptationJingmin Zhu, Anqi Zhu, Hossein Rahmani, Jun Liu et al.NeurIPS 2025 · 3 citations
