Adapting Pre-trained 3D Models for Point Cloud Video Understanding via Cross-frame Spatio-temporal Perception
Baixuan Lv, Yaohua Zha, Tao Dai, Xue Yuerong, Ke Chen, Shu-Tao Xia
Abstract
Point cloud video understanding is becoming increasingly important in fields such as robotics, autonomous driving, and augmented reality, as they can accurately represent object motion and environmental changes. Despite the progress made in self-supervised learning methods for point cloud video understanding, the limited availability of 4D data and the high computational cost of training 4Dspecific models remain significant obstacles. In this paper, we investigate the potential of transferring pre-trained static 3D point cloud models to the 4D domain, pointing out the limitations of static models that capture only spatial information while neglecting temporal dynamics. To address this, we propose a novel Cross-frame Spatio-temporal Adaptation (CSA) strategy by introducing the Point Tube Adapter as the embedding layer and the Geometric Constraint Temporal Adapter (GCTA) to enforce temporal consistency across frames. This strategy extracts both shortterm and long-term temporal dynamics, effectively integrating them with spatial features and enriching the model's understanding of temporal changes in point cloud videos. Extensive experiments on 3D action and gesture recognition tasks demonstrate that our method achieves stateof-the-art performance, establishing its effectiveness for point cloud video understanding.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 58745163-ed37-482a-b294-54c42495ed6eBuilds on23
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point ModelingXumin Yu, Lulu Tang, Yongming Rao, Tiejun Huang et al.CVPR 2022 · 684 citations
- Point-M2AE: Multi-scale Masked Autoencoders for Hierarchical Point Cloud Pre-trainingRenrui Zhang, Ziyu Guo, Peng Gao, Rongyao Fang et al.NeurIPS 2022 · 445 citations
- Self-Supervised Pretraining of 3D Features on any Point-CloudZaiwei Zhang, Rohit Girdhar, Armand Joulin, Ishan MisraICCV 2021 · 333 citations
- Unsupervised Point Cloud Pre-training via Occlusion CompletionHanchen Wang, Qi Liu, Xiangyu Yue, Joan Lasenby et al.ICCV 2021 · 323 citations
Related papers
- X4D-SceneFormer: Enhanced Scene Understanding on 4D Point Cloud Videos through Cross-Modal Knowledge TransferLinglin Jing, Ying Xue, Xu Yan, Chaoda Zheng et al.AAAI 2024 · 14 citations
- Point 4D Transformer Networks for Spatio-Temporal Modeling in Point Cloud VideosHehe Fan, Yi Yang, Mohan S. KankanhalliCVPR 2021
- Mamba4D: Efficient 4D Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space ModelsJiuming Liu, Jinru Han, Lihao Liu, Angelica I. Avilés-Rivero et al.CVPR 2025
- Masked Spatio-Temporal Structure Prediction for Self-supervised Learning on Point Cloud VideosZhiqiang Shen, Xiaoxiao Sheng, Hehe Fan, Longguang Wang et al.ICCV 2023 · 24 citations
- PSTNet: Point Spatio-Temporal Convolution on Point Cloud SequencesHehe Fan, Xin Yu, Yuhang Ding, Yi Yang et al.ICLR 2021 · 148 citations
