Semantic-guided Cross-Modal Prompt Learning for Skeleton-based Zero-shot Action Recognition
Anqi Zhu, Jingmin Zhu, James Bailey, Mingming Gong, Qiuhong Ke
Abstract
Skeleton-based human action recognition is promising due to its privacy preservation, robustness to visual challenges, and computational efficiency. Especially, the practical necessity to recognize unseen actions has led to increased interest in zero-shot skeleton-based action recognition (ZSSAR). Existing ZSSAR approaches often rely on manually crafted action descriptions or visual assumptions to enhance knowledge transfer, which is limited in flexibility and prone to inaccuracies and noise. To overcome this, we introduce Semanticguided Cross-Modal Prompt Learning (SCoPLe), a novel framework that replaces manual guidance with data-driven prompt learning for refinement and alignment of skeletal and textual features. Specifically, we introduce a dual-stream language prompting module that preserves the original semantic context from the pre-trained text encoder while still effectively tuning its ouput for ZSSAR task adaptation. We also introduce a joint-shaped prompting module that learns tuning for skeleton features and incorporate an adaptive visual representation sampler that leverages text semantics to strengthen the cross-modal prompting interactions during skeleton-to-text embedding projection. Experimental results on the NTU-RGB+D and PKU-MMD datasets demonstrate the state-of-the-art performance of our method in both ZS-SAR and generalized ZSSAR scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0aaa9493-1dbc-455c-85aa-5deb97d5f717Cited by top-tier papers5
- Beyond Binary Contrast: Modeling Continuous Skeleton Action Spaces with Transitional AnchorsYingjie Feng, Yi Wang, Jiaze Wang, Anfeng Liu et al.CVPR 2026 · 1 citation
- FineTec: Fine-Grained Action Recognition Under Temporal Corruption via Skeleton Decomposition and Sequence CompletionDian Shao, Mingfei Shi, Like LiuAAAI 2026
- InsAT: Instance-aware Semantic Alignment and Transfer from Human-Object Keypoints for Zero-to-Few-shot Action UnderstandingKazuki TsutsukawaACL 2026
- Universal Skeleton Understanding via Differentiable Rendering and MLLMsZiyi Wang, Peiming Li, Xinshun Wang, Yang Tang et al.ICML 2026
- DarkShake-DVS: Event-based Human Action Recognition under Low-light and Shaking Camera ConditionsJiaqi Chen, Qinfu Xu, Liyuan PanCVPR 2026
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
Related papers
- SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action RecognitionNing Wang, Tieyue Wu, Naeha Sharif, Farid Boussaïd et al.CVPR 2026 · 3 citations
- Fine-Grained Side Information Guided Dual-Prompts for Zero-Shot Skeleton Action RecognitionYang Chen, Jingcai Guo, Tian He, Xiaocheng Lu et al.ACM MM 2024 · 13 citations
- Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-Shot Skeleton-Based Action RecognitionJeonghyeok Do, Munchurl KimICCV 2025 · 6 citations
- Boosting Skeleton-based Zero-Shot Action Recognition with Training-Free Test-Time AdaptationJingmin Zhu, Anqi Zhu, Hossein Rahmani, Jun Liu et al.NeurIPS 2025 · 3 citations
- Generative Action Description Prompts for Skeleton-based Action RecognitionWangmeng Xiang, Chao Li, Yuxuan Zhou, Biao Wang et al.ICCV 2023 · 84 citations
