Exploring Adaptive Masked Reconstruction for Self-Supervised Skeleton-Based Action Recognition
Shengkai Sun, Zhiyong Cheng, Zefan Zhang, Jianfeng Dong, Zhihui Li, Meng Wang
Abstract
Recently, masked skeleton reconstruction models have emerged as strong action representation learners, driving significant progress in self-supervised skeleton‑based action recognition. However, existing state‑of‑the‑art methods must predict an exceedingly large number of spatiotemporal patches, significantly prolonging training time. Besides, by treating all spatiotemporal regions equally during reconstruction, these models are distracted from learning the critical motion patterns that underlie action semantics. To address these challenges, we propose Adaptive Masked Reconstruction (AMR), a faster and stronger pre‑training framework. We first decouple the decoder from the encoder, enabling flexible prediction of larger spatiotemporal patches and dramatically reducing reconstruction complexity. Given that larger patches contain more complex information, which is challenging to predict and consequently degrades performance, we accordingly introduce an adaptive guidance module. This module identifies regions of high motion informativeness, guiding the model to focus on the most discriminative parts of each patch and alleviating reconstruction difficulty. Experiments on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD datasets demonstrate that AMR not only accelerates pre‑training substantially but also improves downstream recognition accuracy, surpassing current state‑of‑the‑art approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 635d2006-02e4-4134-ac92-1d167dfc7b7dBuilds on23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised LearningAdrien Bardes, Jean Ponce, Yann LeCunICLR 2022 · 1,226 citations
- MS2L: Multi-Task Self-Supervised Learning for Skeleton Based Action RecognitionLilang Lin, Sijie Song, Wenhan Yang, Jiaying LiuACM MM 2020 · 217 citations
- Contrastive Learning from Extremely Augmented Skeleton Sequences for Self-Supervised Action RecognitionTianyu Guo, Hong Liu, Zhan Chen, Mengyuan Liu et al.AAAI 2022 · 206 citations
- Skeleton-Contrastive 3D Action Representation LearningFida Mohammad Thoker, Hazel Doughty, Cees G. M. SnoekACM MM 2021 · 158 citations
Related papers
- Towards Efficient General Feature Prediction in Masked Skeleton ModelingShengkai Sun, Zefan Zhang, Jianfeng Dong, Zhiyong Cheng et al.ICCV 2025 · 3 citations
- Masked Motion Predictors are Strong 3D Action Representation LearnersYunyao Mao, Jiajun Deng, Wengang Zhou, Yao Fang et al.ICCV 2023 · 73 citations
- Actionlet-Dependent Contrastive Learning for Unsupervised Skeleton-Based Action RecognitionLilang Lin, Jiahang Zhang, Jiaying LiuCVPR 2023
- SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-trainingHong Yan, Yang Liu, Yushen Wei, Zhen Li et al.ICCV 2023 · 77 citations
- Rethinking Masked Data Reconstruction Pretraining for Strong 3D Action Representation LearningTao Gong, Qi Chu, Bin Liu, Nenghai YuAAAI 2025 · 3 citations
