Masked Motion Predictors are Strong 3D Action Representation Learners
Yunyao Mao, Jiajun Deng, Wengang Zhou, Yao Fang, Wanli Ouyang, Houqiang Li
摘要
In 3D human action recognition, limited supervised data makes it challenging to fully tap into the modeling potential of powerful networks such as transformers. As a result, researchers have been actively investigating effective self-supervised pre-training strategies. In this work, we show that instead of following the prevalent pretext task to perform masked self-component reconstruction in human joints, explicit contextual motion modeling is key to the success of learning effective feature representation for 3D action recognition. Formally, we propose the Masked Motion Prediction (MAMP) framework. To be specific, the proposed MAMP takes as input the masked spatiotemporal skeleton sequence and predicts the corresponding temporal motion of the masked human joints. Considering the high temporal redundancy of the skeleton sequence, in our MAMP, the motion information also acts as an empirical semantic richness prior that guide the masking process, promoting better attention to semantically rich temporal regions. Extensive experiments on NTU-60, NTU-120, and PKU-MMD datasets show that the proposed MAMP pre-training substantially improves the performance of the adopted vanilla transformer, achieving state-of-the-art results without bells and whistles. The source code of our MAMP is available at https:// github.com/maoyunyao/MAMP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- USDRL: Unified Skeleton-Based Dense Representation Learning with Multi-Grained Feature DecorrelationWanjiang Weng, Hongsong Wang, Junbo Wang, Lei He 等AAAI 2025 · 被引用 15 次
- Masked Pre-training Enables Universal Zero-shot DenoiserXiaoxiao Ma, Zhixiang Wei, Yi Jin, Pengyang Ling 等NeurIPS 2024 · 被引用 11 次
- CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action RecognitionYuhang Wen, Mengyuan Liu, Songtao Wu, Beichen DingNeurIPS 2024 · 被引用 7 次
- Learning in Order! A Sequential Strategy to Learn Invariant Features for Multimodal Sentiment AnalysisXianbing Zhao, Lizhen Qu, Tao Feng, Jianfei Cai 等ACM MM 2024 · 被引用 3 次
- Towards Efficient General Feature Prediction in Masked Skeleton ModelingShengkai Sun, Zefan Zhang, Jianfeng Dong, Zhiyong Cheng 等ICCV 2025 · 被引用 3 次
它引用的顶会 Paper23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionYuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li 等ICCV 2021 · 被引用 871 次
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin 等CVPR 2022 · 被引用 752 次
相关 Paper
- Rethinking Masked Data Reconstruction Pretraining for Strong 3D Action Representation LearningTao Gong, Qi Chu, Bin Liu, Nenghai YuAAAI 2025 · 被引用 3 次
- Prompted Contrast with Masked Motion Modeling: Towards Versatile 3D Action Representation LearningJiahang Zhang, Lilang Lin, Jiaying LiuACM MM 2023 · 被引用 26 次
- Exploring Adaptive Masked Reconstruction for Self-Supervised Skeleton-Based Action RecognitionShengkai Sun, Zhiyong Cheng, Zefan Zhang, Jianfeng Dong 等CVPR 2026 · 被引用 2 次
- Self-Supervised Action Representation Learning from Partial Spatio-Temporal Skeleton SequencesYujie Zhou, Haodong Duan, Anyi Rao, Bing Su 等AAAI 2023 · 被引用 62 次
- A Dual-Masked Auto-Encoder for Robust Motion Capture with Spatial-Temporal Skeletal Token CompletionJunkun Jiang, Jie Chen, Yike GuoACM MM 2022 · 被引用 8 次
