SkeletonMAE: Graph-based Masked Autoencoder for Skeleton Sequence Pre-training
Hong Yan, Yang Liu, Yushen Wei, Zhen Li, Guanbin Li, Liang Lin
摘要
Skeleton sequence representation learning has shown great advantages for action recognition due to its promising ability to model human joints and topology. However, the current methods usually require sufficient labeled data for training computationally expensive models, which is labor-intensive and time-consuming. Moreover, these methods ignore how to utilize the fine-grained dependencies among different skeleton joints to pre-train an efficient skeleton sequence learning model that can generalize well across different datasets. In this paper, we propose an efficient skeleton sequence learning framework, named Skeleton Sequence Learning (SSL). To comprehensively capture the human pose and obtain discriminative skeleton sequence representation, we build an asymmetric graph-based encoder-decoder pre-training architecture named SkeletonMAE, which embeds skeleton joint sequence into Graph Convolutional Network (GCN) and reconstructs the masked skeleton joints and edges based on the prior human topology knowledge. Then, the pre-trained Skele-tonMAE encoder is integrated with the Spatial-Temporal Representation Learning (STRL) module to build the SSL framework. Extensive experimental results show that our SSL generalizes well across different datasets and outperforms the state-of-the-art self-supervised skeleton-based action recognition methods on FineGym, Diving48, NTU 60 and NTU 120 datasets. Additionally, we obtain comparable performance to some fully supervised methods. The code is avaliable at https://github.com/HongYan1123/ SkeletonMAE
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- A Dual-Augmentor Framework for Domain Generalization in 3D Human Pose EstimationQucheng Peng, Ce Zheng, Chen ChenCVPR 2024 · 被引用 38 次
- USDRL: Unified Skeleton-Based Dense Representation Learning with Multi-Grained Feature DecorrelationWanjiang Weng, Hongsong Wang, Junbo Wang, Lei He 等AAAI 2025 · 被引用 15 次
- Toward Approaches to Scalability in 3D Human Pose EstimationJun-Hui Kim, Seong-Whan LeeNeurIPS 2024 · 被引用 5 次
- 3DAffordSplat: Efficient Affordance Reasoning with 3D GaussiansZeming Wei, Junyi Lin, Yang Liu, Weixing Chen 等ACM MM 2025 · 被引用 4 次
- Hi-GMAE: Hierarchical Graph Masked AutoencodersChuang Liu, Zelin Yao, Xueqi Ma, Mukun Chen 等WWW 2026 · 被引用 3 次
它引用的顶会 Paper32
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 被引用 4,104 次
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-TrainingZhan Tong, Yibing Song, Jue Wang, Limin WangNeurIPS 2022 · 被引用 2,336 次
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionYuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li 等ICCV 2021 · 被引用 871 次
相关 Paper
- Dynamic Semantic-Based Spatial Graph Convolution Network for Skeleton-Based Human Action RecognitionJianyang Xie, Yanda Meng, Yitian Zhao, Anh Nguyen 等AAAI 2024 · 被引用 59 次
- Graph Contrastive Learning for Skeleton-based Action RecognitionXiaohu Huang, Hao Zhou, Jian Wang, Haocheng Feng 等ICLR 2023 · 被引用 13 次
- Towards Efficient General Feature Prediction in Masked Skeleton ModelingShengkai Sun, Zefan Zhang, Jianfeng Dong, Zhiyong Cheng 等ICCV 2025 · 被引用 3 次
- Skeleton-Contrastive 3D Action Representation LearningFida Mohammad Thoker, Hazel Doughty, Cees G. M. SnoekACM MM 2021 · 被引用 158 次
- SkeleTR: Towards Skeleton-based Action Recognition in the WildHaodong Duan, Mingze Xu, Bing Shuai, Davide Modolo 等ICCV 2023 · 被引用 38 次
