Prompted Contrast with Masked Motion Modeling: Towards Versatile 3D Action Representation Learning
Jiahang Zhang, Lilang Lin, Jiaying Liu
Abstract
Self-supervised learning has proved effective for skeleton-based human action understanding, which is an important yet challenging topic. Previous works mainly rely on contrastive learning or masked motion modeling paradigm to model the skeleton relations. However, the sequence-level and joint-level representation learning cannot be effectively and simultaneously handled by these methods. As a result, the learned representations fail to generalize to different downstream tasks. Moreover, combining these two paradigms in a naive manner leaves the synergy between them untapped and can lead to interference in training. To address these problems, we propose Prompted Contrast with Masked Motion Modeling, PCM 3 , for versatile 3D action representation learning. Our method integrates the contrastive learning and masked prediction tasks in a mutually beneficial manner, which substantially boosts the generalization capacity for various downstream tasks. Specifically, masked prediction provides novel training views for contrastive learning, which in turn guides the masked prediction training with high-level semantic information. Moreover, we propose a dualprompted multi-task pretraining strategy, which further improves model representations by reducing the interference caused by learning the two different pretext tasks. Extensive experiments on five downstream tasks under three large-scale datasets are conducted, demonstrating the superior generalization capacity of PCM 3 compared to the state-of-the-art works. Our project is publicly available at: https://jhang2020.github.io/Projects/PCM3/PCM3.html.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6530e997-6cce-4092-b57d-0a21997fd17dCited by top-tier papers8
- USDRL: Unified Skeleton-Based Dense Representation Learning with Multi-Grained Feature DecorrelationWanjiang Weng, Hongsong Wang, Junbo Wang, Lei He et al.AAAI 2025 · 15 citations
- Towards Efficient General Feature Prediction in Masked Skeleton ModelingShengkai Sun, Zefan Zhang, Jianfeng Dong, Zhiyong Cheng et al.ICCV 2025 · 3 citations
- Exploring Adaptive Masked Reconstruction for Self-Supervised Skeleton-Based Action RecognitionShengkai Sun, Zhiyong Cheng, Zefan Zhang, Jianfeng Dong et al.CVPR 2026 · 2 citations
- Beyond Binary Contrast: Modeling Continuous Skeleton Action Spaces with Transitional AnchorsYingjie Feng, Yi Wang, Jiaze Wang, Anfeng Liu et al.CVPR 2026 · 1 citation
- Action Motifs: Self-Supervised Hierarchical Representation of Human Body MovementsGenki Kinoshita, Shu Nakamura, Ryo Kawahara, Shohei Nobuhara et al.CVPR 2026
Builds on21
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin et al.CVPR 2022 · 1,129 citations
- MS2L: Multi-Task Self-Supervised Learning for Skeleton Based Action RecognitionLilang Lin, Sijie Song, Wenhan Yang, Jiaying LiuACM MM 2020 · 217 citations
- Contrast with Reconstruct: Contrastive 3D Representation Learning Guided by Generative PretrainingZekun Qi, Runpei Dong, Guofan Fan, Zheng Ge et al.ICML 2023 · 209 citations
Related papers
- Rethinking Masked Data Reconstruction Pretraining for Strong 3D Action Representation LearningTao Gong, Qi Chu, Bin Liu, Nenghai YuAAAI 2025 · 3 citations
- Modeling the Uncertainty for Self-supervised 3D Skeleton Action Representation LearningYukun Su, Guosheng Lin, Ruizhou Sun, Yun Hao et al.ACM MM 2021 · 30 citations
- Masked Motion Predictors are Strong 3D Action Representation LearnersYunyao Mao, Jiajun Deng, Wengang Zhou, Yao Fang et al.ICCV 2023 · 73 citations
- Skeleton-Contrastive 3D Action Representation LearningFida Mohammad Thoker, Hazel Doughty, Cees G. M. SnoekACM MM 2021 · 158 citations
- LiftedCL: Lifting Contrastive Learning for Human-Centric PerceptionZiwei Chen, Qiang Li, Xiaofeng Wang, Wankou YangICLR 2023
