Zero-shot Skeleton-based Action Recognition via Mutual Information Estimation and Maximization
Yujie Zhou, Wenwen Qiang, Anyi Rao, Ning Lin, Bing Su, Jiaqi Wang
Abstract
Zero-shot skeleton-based action recognition aims to recognize actions of unseen categories after training on data of seen categories.
The key is to build the connection between visual and semantic space from seen to unseen classes. Previous studies have primarily focused on encoding sequences into a singular feature vector, with subsequent mapping the features to an identical anchor point within the embedded space. Their performance is hindered by 1) the ignorance of the global visual/semantic distribution alignment, which results in a limitation to capture the true interdependence between the two spaces. 2) the negligence of temporal information since the frame-wise features with rich action clues are directly pooled into a single feature vector. We propose a new zero-shot skeleton-based action recognition method via mutual information (MI) estimation and maximization. Specifically, 1) we maximize the MI between visual and semantic space for distribution alignment; 2) we leverage the temporal information for estimating the MI by encouraging MI to increase as more frames are observed. Extensive experiments on three large-scale skeleton action datasets confirm the effectiveness of our method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Part-Aware Unified Representation of Language and Skeleton for Zero-Shot Action RecognitionAnqi Zhu, Qiuhong Ke, Mingming Gong, James BaileyCVPR 2024 · 16 citations
- Fine-Grained Side Information Guided Dual-Prompts for Zero-Shot Skeleton Action RecognitionYang Chen, Jingcai Guo, Tian He, Xiaocheng Lu et al.ACM MM 2024 · 13 citations
- Bridging the Skeleton-Text Modality Gap: Diffusion-Powered Modality Alignment for Zero-Shot Skeleton-Based Action RecognitionJeonghyeok Do, Munchurl KimICCV 2025 · 6 citations
- SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily LivingArkaprava Sinha, Dominick Reilly, François Brémond, Pu Wang et al.AAAI 2025 · 5 citations
- Frequency-Semantic Enhanced Variational Autoencoder for Zero-Shot Skeleton-Based Action RecognitionWenhan Wu, Zhishuai Guo, Chen Chen, Hongfei Xue et al.ICCV 2025 · 4 citations
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- On Mutual Information Maximization for Representation LearningMichael Tschannen, Josip Djolonga, Paul K. Rubenstein, Sylvain Gelly et al.ICLR 2020 · 559 citations
- MS2L: Multi-Task Self-Supervised Learning for Skeleton Based Action RecognitionLilang Lin, Sijie Song, Wenhan Yang, Jiaying LiuACM MM 2020 · 217 citations
- Contrastive Learning from Extremely Augmented Skeleton Sequences for Self-Supervised Action RecognitionTianyu Guo, Hong Liu, Zhan Chen, Mengyuan Liu et al.AAAI 2022 · 206 citations
- Skeleton-Contrastive 3D Action Representation LearningFida Mohammad Thoker, Hazel Doughty, Cees G. M. SnoekACM MM 2021 · 158 citations
Related papers
- Neuron: Learning Context-Aware Evolving Representations for Zero-Shot Skeleton Action RecognitionYang Chen, Jingcai Guo, Song Guo, Dacheng TaoCVPR 2025
- Zero-Shot Learning for IMU-Based Activity Recognition Using Video EmbeddingsCatherine Tong, Jinchen Ge, Nicholas D. LaneUbiComp 2022 · 39 citations
- SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action RecognitionNing Wang, Tieyue Wu, Naeha Sharif, Farid Boussaïd et al.CVPR 2026 · 3 citations
- Semantic-guided Cross-Modal Prompt Learning for Skeleton-based Zero-shot Action RecognitionAnqi Zhu, Jingmin Zhu, James Bailey, Mingming Gong et al.CVPR 2025
- Boosting Skeleton-based Zero-Shot Action Recognition with Training-Free Test-Time AdaptationJingmin Zhu, Anqi Zhu, Hossein Rahmani, Jun Liu et al.NeurIPS 2025 · 3 citations
