Masked Motion Predictors are Strong 3D Action Representation Learners
Yunyao Mao, Jiajun Deng, Wengang Zhou, Yao Fang, Wanli Ouyang, Houqiang Li
Abstract
In 3D human action recognition, limited supervised data makes it challenging to fully tap into the modeling potential of powerful networks such as transformers. As a result, researchers have been actively investigating effective self-supervised pre-training strategies. In this work, we show that instead of following the prevalent pretext task to perform masked self-component reconstruction in human joints, explicit contextual motion modeling is key to the success of learning effective feature representation for 3D action recognition. Formally, we propose the Masked Motion Prediction (MAMP) framework. To be specific, the proposed MAMP takes as input the masked spatiotemporal skeleton sequence and predicts the corresponding temporal motion of the masked human joints. Considering the high temporal redundancy of the skeleton sequence, in our MAMP, the motion information also acts as an empirical semantic richness prior that guide the masking process, promoting better attention to semantically rich temporal regions. Extensive experiments on NTU-60, NTU-120, and PKU-MMD datasets show that the proposed MAMP pre-training substantially improves the performance of the adopted vanilla transformer, achieving state-of-the-art results without bells and whistles. The source code of our MAMP is available at https:// github.com/maoyunyao/MAMP .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 87bfe73a-c3a4-4497-8546-ebb5fc185c95Cited by top-tier papers14
- USDRL: Unified Skeleton-Based Dense Representation Learning with Multi-Grained Feature DecorrelationWanjiang Weng, Hongsong Wang, Junbo Wang, Lei He et al.AAAI 2025 · 15 citations
- Masked Pre-training Enables Universal Zero-shot DenoiserXiaoxiao Ma, Zhixiang Wei, Yi Jin, Pengyang Ling et al.NeurIPS 2024 · 11 citations
- CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action RecognitionYuhang Wen, Mengyuan Liu, Songtao Wu, Beichen DingNeurIPS 2024 · 7 citations
- Learning in Order! A Sequential Strategy to Learn Invariant Features for Multimodal Sentiment AnalysisXianbing Zhao, Lizhen Qu, Tao Feng, Jianfei Cai et al.ACM MM 2024 · 3 citations
- Towards Efficient General Feature Prediction in Masked Skeleton ModelingShengkai Sun, Zefan Zhang, Jianfeng Dong, Zhiyong Cheng et al.ICCV 2025 · 3 citations
Builds on23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionYuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li et al.ICCV 2021 · 871 citations
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin et al.CVPR 2022 · 752 citations
Related papers
- Rethinking Masked Data Reconstruction Pretraining for Strong 3D Action Representation LearningTao Gong, Qi Chu, Bin Liu, Nenghai YuAAAI 2025 · 3 citations
- Prompted Contrast with Masked Motion Modeling: Towards Versatile 3D Action Representation LearningJiahang Zhang, Lilang Lin, Jiaying LiuACM MM 2023 · 26 citations
- Exploring Adaptive Masked Reconstruction for Self-Supervised Skeleton-Based Action RecognitionShengkai Sun, Zhiyong Cheng, Zefan Zhang, Jianfeng Dong et al.CVPR 2026 · 2 citations
- Self-Supervised Action Representation Learning from Partial Spatio-Temporal Skeleton SequencesYujie Zhou, Haodong Duan, Anyi Rao, Bing Su et al.AAAI 2023 · 62 citations
- A Dual-Masked Auto-Encoder for Robust Motion Capture with Spatial-Temporal Skeletal Token CompletionJunkun Jiang, Jie Chen, Yike GuoACM MM 2022 · 8 citations
