USDRL: Unified Skeleton-Based Dense Representation Learning with Multi-Grained Feature Decorrelation
Wanjiang Weng, Hongsong Wang, Junbo Wang, Lei He, Guo-Sen Xie
Abstract
Contrastive learning has achieved great success in skeletonbased representation learning recently. However, the prevailing methods are predominantly negative-based, necessitating additional momentum encoder and memory bank to get negative samples, which increases the difficulty of model training. Furthermore, these methods primarily concentrate on learning a global representation for recognition and retrieval tasks, while overlooking the rich and detailed local representations that are crucial for dense prediction tasks. To alleviate these issues, we introduce a Unified Skeleton-based Dense Representation Learning framework based on feature decorrelation, called USDRL, which employs feature decorrelation across temporal, spatial, and instance domains in a multi-grained manner to reduce redundancy among dimensions of the representations to maximize information extraction from features. Additionally, we design a Dense Spatio-Temporal Encoder (DSTE) to capture fine-grained action representations effectively, thereby enhancing the performance of dense prediction tasks. Comprehensive experiments, conducted on the benchmarks NTU-60, NTU-120, PKU-MMD I, and PKU-MMD II, across diverse downstream tasks including action recognition, action retrieval, and action detection, conclusively demonstrate that our approach significantly outperforms the current state-of-the-art (SOTA) approaches. Our code and models are available at https://github.com/wengwanjiang/USDRL .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2823f12-e01b-4f47-aad2-4cd3061f5099Cited by top-tier papers7
- Dual Conditioned Motion Diffusion for Pose-Based Video Anomaly DetectionHongsong Wang, Andi Xu, Pinle Ding, Jie GuiAAAI 2025 · 8 citations
- Towards Efficient General Feature Prediction in Masked Skeleton ModelingShengkai Sun, Zefan Zhang, Jianfeng Dong, Zhiyong Cheng et al.ICCV 2025 · 3 citations
- Exploring Adaptive Masked Reconstruction for Self-Supervised Skeleton-Based Action RecognitionShengkai Sun, Zhiyong Cheng, Zefan Zhang, Jianfeng Dong et al.CVPR 2026 · 2 citations
- Beyond Binary Contrast: Modeling Continuous Skeleton Action Spaces with Transitional AnchorsYingjie Feng, Yi Wang, Jiaze Wang, Anfeng Liu et al.CVPR 2026 · 1 citation
- Action Motifs: Self-Supervised Hierarchical Representation of Human Body MovementsGenki Kinoshita, Shu Nakamura, Ryo Kawahara, Shohei Nobuhara et al.CVPR 2026
Builds on28
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised LearningAdrien Bardes, Jean Ponce, Yann LeCunICLR 2022 · 1,226 citations
- SimMIM: a Simple Framework for Masked Image ModelingZhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin et al.CVPR 2022 · 1,129 citations
- Understanding Dimensional Collapse in Contrastive Self-supervised LearningLi Jing, Pascal Vincent, Yann LeCun, Yuandong TianICLR 2022 · 467 citations
Related papers
- SCD-Net: Spatiotemporal Clues Disentanglement Network for Self-Supervised Skeleton-Based Action RecognitionCong Wu, Xiao-Jun Wu, Josef Kittler, Tianyang Xu et al.AAAI 2024 · 29 citations
- Skeleton-Contrastive 3D Action Representation LearningFida Mohammad Thoker, Hazel Doughty, Cees G. M. SnoekACM MM 2021 · 158 citations
- Hierarchical Contrast for Unsupervised Skeleton-Based Action Representation LearningJianfeng Dong, Shengkai Sun, Zhonglin Liu, Shujie Chen et al.AAAI 2023 · 73 citations
- Modeling the Relative Visual Tempo for Self-supervised Skeleton-based Action RecognitionYisheng Zhu, Hu Han, Zhengtao Yu, Guangcan LiuICCV 2023 · 32 citations
- Self-Supervised Action Representation Learning from Partial Spatio-Temporal Skeleton SequencesYujie Zhou, Haodong Duan, Anyi Rao, Bing Su et al.AAAI 2023 · 62 citations
