Point Contrastive Prediction with Semantic Clustering for Self-Supervised Learning on Point Cloud Videos
Xiaoxiao Sheng, Zhiqiang Shen, Gang Xiao, Longguang Wang, Yulan Guo, Hehe Fan
Abstract
We propose a unified point cloud video self-supervised learning framework for object-centric and scene-centric data. Previous methods commonly conduct representation learning at the clip or frame level and cannot well capture fine-grained semantics. Instead of contrasting the representations of clips or frames, in this paper, we propose a unified self-supervised framework by conducting contrastive learning at the point level. Moreover, we introduce a new pretext task by achieving semantic alignment of superpoints, which further facilitates the representations to capture semantic cues at multiple scales. In addition, due to the high redundancy in the temporal dimension of dynamic point clouds, directly conducting contrastive learning at the point level usually leads to massive undesired negatives and insufficient modeling of positive representations. To remedy this, we propose a selection strategy to retain proper negatives and make use of high-similarity samples from other instances as positive supplements. Extensive experiments show that our method outperforms supervised counterparts on a wide range of downstream tasks and demonstrates the superior transferability of the learned representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 01e27d57-ff71-4100-9275-2a8284ca9824Cited by top-tier papers7
- A Unified Framework for Human-centric Point Cloud Video UnderstandingYiteng Xu, Kecheng Ye, Xiao Han, Yiming Ren et al.CVPR 2024 · 4 citations
- UST-SSM: Unified Spatio-Temporal State Space Models for Point Cloud Video ModelingPeiming Li, Ziyi Wang, Yulin Yuan, Hong Liu et al.ICCV 2025 · 3 citations
- Recognizing Actions From Robotic View for Natural Human-Robot InteractionZiyi Wang, Peiming Li, Hong Liu, Zhichao Deng et al.ICCV 2025 · 1 citation
- DeSPITE: Exploring Contrastive Deep Skeleton-Pointcloud-IMU-Text Embeddings for Advanced Point Cloud Human Activity UnderstandingThomas Kreutz, Max Mühlhäuser, Alejandro Sánchez GuineaICCV 2025 · 1 citation
- Adapting Pre-trained 3D Models for Point Cloud Video Understanding via Cross-frame Spatio-temporal PerceptionBaixuan Lv, Yaohua Zha, Tao Dai, Xue Yuerong et al.CVPR 2025
Builds on30
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
Related papers
- PointCMP: Contrastive Mask Prediction for Self-supervised Learning on Point Cloud VideosZhiqiang Shen, Xiaoxiao Sheng, Longguang Wang, Yulan Guo et al.CVPR 2023
- GroupContrast: Semantic-Aware Self-Supervised Representation Learning for 3D UnderstandingChengyao Wang, Li Jiang, Xiaoyang Wu, Zhuotao Tian et al.CVPR 2024 · 18 citations
- ToThePoint: Efficient Contrastive Learning of 3D Point Clouds via RecyclingXinglin Li, Jiajing Chen, Jinhui Ouyang, Hanhui Deng et al.CVPR 2023
- Point-GCC: Universal Self-supervised 3D Scene Pre-training via Geometry-Color ContrastGuofan Fan, Zekun Qi, Wenkai Shi, Kaisheng MaACM MM 2024 · 12 citations
- Self-Contrastive Learning with Hard Negative Sampling for Self-supervised Point Cloud LearningBi'an Du, Xiang Gao, Wei Hu, Xin LiACM MM 2021 · 82 citations
